Architecture Note: Designing Boundaries and Failure Isolation in a Network Policy Service
An architecture note on isolating policy decisions from data-plane forwarding in a network policy service, with modular components, trust boundaries, health checks, and failure-mode mitigations.
03 Jun 2026, 21:07 UTC

The Problem: Policy Enforcement Without Cascading Failure
Network policy services sit in the critical path of every workload connection. When the policy layer becomes unavailable or slow, the impact radiates outward—applications time out, retries amplify load, and operators lose visibility into whether the network or the policy engine is at fault. This note describes a minimal architecture that isolates policy decisions from data-plane forwarding, defines clear trust boundaries, and provides operational signals that distinguish healthy degradation from systemic failure.
Requirements Driving the Design
- Decision latency budget: Policy evaluation must complete within a configurable SLA (e.g., p99 < 5 ms) so that sidecar proxies or CNI plugins can enforce without adding perceptible latency.
- Independent scaling: Authentication, authorization, and policy compilation workloads have different traffic profiles and must scale separately.
- Auditability: Every allow/deny decision must be traceable to a specific policy version and identity context.
- Graceful degradation: If the policy engine is unreachable, the data plane should fail closed (deny) or fail open (allow) based on a declared posture, not an implicit timeout.
Smallest Suitable Design: Three Deployable Modules
Instead of a monolithic policy server, split into three independently deployable services that communicate over gRPC with explicit contracts:
| Module | Responsibility | Scaling Trigger |
|---|---|---|
npss-auth | Validates caller identity (mTLS, JWT, SPIFFE) and emits a normalized principal | Request rate, certificate rotation events |
npss-policy | Compiles YAML/CRD policies into an internal decision DAG; serves evaluation requests | Policy change rate, evaluation latency |
npss-store | Persists policy versions, audit logs, and principal-to-policy mappings; provides read replicas | Write throughput, replication lag |
Each module exposes its own /healthz (liveness) and /readyz (readiness) endpoints. The readiness check for npss-policy verifies it can reach a quorum of npss-store replicas and that its local policy cache is within maxStalenessSeconds of the latest committed version.
Trust and Data Boundaries
Draw hard lines between three zones:
- Ingress zone: Accepts external API calls (e.g., from a CNI plugin). Terminates mTLS, validates the caller against a allowlist of trusted sidecar identities. No policy logic runs here.
- Evaluation zone: Stateless
npss-policyreplicas. They receive a normalized principal and a connection tuple (src/dst IP, port, protocol). They returnALLOW,DENY, orERRORwith a policy version identifier. They never write to persistent storage. - Storage zone:
npss-storecluster (e.g., etcd or PostgreSQL with synchronous replication). Only the control-plane controller andnpss-policyon cache miss may read; only the controller writes.
Data crossing boundaries is versioned: every evaluation response includes policyVersion: "v2024.03.15-14-gabc123". The caller can detect stale decisions by comparing this version against the version advertised in npss-store's /readyz payload.
Operational Checks That Distinguish Degradation from Outage
Deploy the following checks as part of your service mesh or Kubernetes probes:
# Example readiness probe for npss-policy (run inside the pod)
curl -sf http://localhost:8080/readyz | jq -e '.storeReachable == true and .cacheStalenessSec < 30'
Where to run: Kubernetes kubelet executes this inside the npss-policy container. Requires curl and jq in the image.
Expected check: Returns HTTP 200 only when the module can reach a quorum of store replicas and its local policy cache is fresh. If the store is reachable but cache staleness exceeds 30 seconds, the pod is marked not ready—traffic shifts to healthier replicas.
Risk: Setting maxStalenessSeconds too low causes flapping during rolling updates; too high serves stale policy. Start with 30 seconds and tune based on observed policy propagation latency.
Failure Modes and Mitigations
Cascading Timeouts
A slow npss-store read causes npss-policy to block, which causes sidecar proxies to queue requests, which triggers retries that amplify load on the store.
Mitigation: Configure a circuit breaker in the gRPC client from npss-policy to npss-store:
# npss-policy config snippet
circuitBreaker:
maxConcurrentCalls: 500
timeout: 200ms
errorThresholdPercentage: 50
sleepWindow: 10s
When the breaker trips, npss-policy serves from its local cache (if within staleness bound) or returns ERROR with reason: "STORE_CIRCUIT_OPEN". The sidecar proxy treats ERROR according to its configured fail posture.
Stale Cache Serving Revoked Policy
If npss-store accepts a policy revocation but npss-policy replicas haven't refreshed, they may allow traffic that should be denied.
Mitigation: Use a lease-based invalidation. On policy write, the controller writes a policyVersion key with a TTL equal to maxStalenessSeconds. npss-policy watches this key; expiry forces a synchronous cache refresh before serving new requests. If the watch disconnects, the module enters a DEGRADED readiness state.
Split-Brain in Storage Zone
Network partition isolates a minority of npss-store replicas. The majority continues serving reads; the minority must reject writes and signal unavailability.
Mitigation: Run npss-store with a consensus protocol (Raft) that requires quorum for both reads and writes. Configure npss-policy readiness to require storeQuorumHealthy: true.
Conditions That Would Change This Design
- Sub-millisecond latency requirement: Move policy evaluation into the data plane (e.g., eBPF in the CNI plugin) and use the service only for compilation and distribution.
- Multi-tenancy with hard isolation: Deploy separate
npss-policy+npss-storestacks per tenant; share onlynpss-auth. - Policy language requires global analysis (e.g., conflict detection across namespaces): Add a
npss-compilermodule that runs as a batch job before publishing tonpss-store. - Audit log volume exceeds storage zone capacity: Offload audit writes to an append-only log (Kafka, Cloud Pub/Sub) and keep
npss-store for current state only.
Limitations and Verification Steps
This architecture assumes a generic Npss (Network Policy Service Surface) model. Actual implementations—whether Calico, Cilium, Istio AuthorizationPolicy, or a vendor appliance—vary in:
- Supported health-check endpoints and response schemas
- Whether policy compilation is incremental or full-rebuild
- Default fail posture (closed vs. open) and its configurability
Verify before adopting:
- Check your Npss documentation for
/readyzand/healthzschemas. Confirm they expose store connectivity and cache staleness. - Run integration tests that:
- Deploy the three modules with induced network latency between policy and store.
- Measure evaluation latency percentiles under load.
- Simulate store partition and verify circuit breaker trips within configured timeout.
- Validate that stale-cache detection triggers readiness failure.
- Confirm the version identifier format matches your GitOps or CI/CD pipeline so that audit traces are actionable.
If your Npss release does not expose cache staleness or store quorum status in readiness, the three-module split adds operational complexity without the intended isolation. In that case, keep a single deployment but enforce the same logical boundaries via internal package structure and distinct configuration sections.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.