Architecture Note: Implementing ASP.NET Core Health Checks in .NET Services
An architecture‑focused guide for adding ASP.NET Core Health Checks: requirements, minimal design, trust boundaries, operational settings, failure modes, and triggers for redesign.
08 Jan 2026, 06:40 UTC

Requirements
Before adding health checks, clarify what the service must expose:
- An HTTP endpoint that returns
200 OKwhen all registered checks pass and503 Service Unavailableotherwise. - Payload that indicates overall status and, optionally, per‑check details without leaking secrets.
- Ability to configure timeouts, latency thresholds, and result caching to avoid thundering‑herd scrapes.
- Clear trust boundary: checks run in‑process with the same privileges as the host, so any diagnostic data must be sanitized before being returned.
Minimal Design
The smallest viable implementation uses the built‑in HealthCheckService and the default JSON writer. Register only the checks that are essential for determining service readiness (e.g., database connectivity, critical external API, disk space).
Example: Minimal ASP.NET Core Web API
// Program.cs (ASP.NET Core 6+ minimal hosting model)
var builder = WebApplication.CreateBuilder(args);
// Add health checks services
builder.Services.AddHealthChecks()
// Example: a lightweight lambda check
.AddCheck("self", () => HealthCheckResult.Healthy());
var app = builder.Build();
// Map the health endpoint at /health
app.MapHealthChecks("/health");
app.Run();
To try this locally:
- Open a terminal in an empty folder.
- Run
dotnet new web -n HealthDemoto create a project. - Change into the folder:
cd HealthDemo. - Add the health checks package:
dotnet add package Microsoft.AspNetCore.Diagnostics.HealthChecks. - Replace the generated
Program.cswith the snippet above. - Execute
dotnet run(requires .NET SDK 6.0 or later). - In another terminal, call
curl -i http://localhost:5000/health(or use a browser). You should see a200response with JSON similar to {"status":"Healthy"}.
Trust/Data Boundaries
Health check callbacks execute with the same process identity as the host. Avoid returning connection strings, passwords, or any internal tokens in the check result. If a check needs to inspect sensitive data, sanitize it before exposing it in the HealthCheckResult description or data dictionary.
Custom ResponseWriter Example
// In Program.cs, after builder.Services.AddHealthChecks()
builder.Services.AddHealthChecks()
.AddCheck("db", () => {
// Imagine a real DB ping here; do not return the connection string.
return HealthCheckResult.Healthy();
});
var app = builder.Build();
app.MapHealthChecks("/health", new HealthCheckOptions
{
// Omit the full JSON payload; write a minimal safe version.
ResponseWriter = async (c, r) =>
{
var result = JsonSerializer.Serialize(new
{
status = r.Status.ToString(),
// Only include non‑sensitive check names
checks = r.Entries.Select(e => new { name = e.Key, status = e.Value.Status.ToString() })
});
c.Response.ContentType = "application/json";
await c.Response.WriteAsync(result);
}
});
app.Run();
After deploying this change, a request to /health will return only the status and check names, ensuring no connection strings appear in the payload.
Operational Checks
To keep health checks lightweight and prevent them from degrading request latency:
- Set a timeout per registration:
.AddCheck("api", ..., TimeSpan.FromSeconds(2)). - Define a latency threshold via
HealthCheckOptions.AllowCachingResponsesandCacheDurationso repeated scrapes within a short window reuse the last result. - Offload expensive work (e.g., large file scans) to a background timer and have the health check read a pre‑computed flag.
Configuration Example with Caching
builder.Services.AddHealthChecks()
.AddCheck("disk", () => {
// Simulate a quick disk‑space check
var free = new DriveInfo("C").AvailableFreeSpace;
return free > 10L * 1024 * 1024 * 1024 ? HealthCheckResult.Healthy() : HealthCheckResult.Degraded();
});
var app = builder.Build();
app.MapHealthChecks("/health", new HealthCheckOptions
{
// Cache the result for 10 seconds to reduce load during frequent scrapes
AllowCachingResponses = true,
CacheDuration = TimeSpan.FromSeconds(10)
});
app.Run();
With caching enabled, successive GET requests to /health within the cache window will return the same payload without re‑executing the disk check.
Failure Modes
Health checks can fail in several ways:
- Unhandled exception inside a check → the middleware treats the check as
Unhealthyand includes the exception message in the result (ensure messages are sanitized). - Timeout expiration → the check is marked
Unhealthywith a timeout descriptor. - Resource exhaustion (e.g., thread pool starvation) if checks are long‑running; this can increase overall latency and cause cascading failures.
Monitoring the health endpoint itself (e.g., via an external alerting system) helps detect when the service is unable to respond due to these failures.
When the Design Should Change
Revisit the health‑check architecture if any of the following conditions arise:
- New critical dependencies are introduced (e.g., a message broker, cache, or third‑party SDK) that must be reflected in readiness decisions.
- Operational data shows that the current check set is causing noticeable latency spikes during peak scrape intervals.
- Security reviews reveal that diagnostic data is being leaked; then introduce a stricter
ResponseWriteror move sensitive checks to a sidecar process. - The service migrates to a different hosting model (e.g., Azure Functions, worker service) where the built‑in middleware is unavailable; a custom health‑check implementation may be required.
Limitations and Practical Verification
This note describes the built‑in ASP.NET Core health‑check framework; it does not cover third‑party alternatives (e.g., AspNetCore.HealthChecks.SqlServer, AspNetCore.HealthChecks.UI). Verify that your chosen checks meet the latency and safety requirements of your production environment by:
- Deploying the service to a staging environment with realistic load.
- Scraping the
/healthendpoint at the frequency intended for production (e.g., every 15 seconds) and measuring response time. - Introducing a failing check (return
HealthCheckResult.Unhealthy()) and confirming the endpoint returns503with the expected JSON payload. - Reviewing the payload for any accidental leakage of secrets; use a tool like
grepon the response to search for known patterns (connection strings, passwords).
If any of these verification steps reveal excessive latency, timeouts, or data exposure, adjust the check implementation, timeout values, or caching strategy accordingly.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.