Automated Cohort Identification in Clarity: When Configurable Clinical Criteria Meet Messy Real-World Data
Clarity's automated cohort identification combines configurable clinical criteria with NLP over clinical notes. This blog walks through the engineering decisions—evidence weighting, FHIR data readiness rules, audit trails—and the trade-offs around vendor lock-in and model drift.
08 Oct 2025, 23:17 UTC

The Problem: Manual Chart Review Doesn't Scale
Every healthcare analytics team hits the same wall: identifying patients for a quality measure, clinical trial, or population health intervention requires someone to write SQL against a normalized EHR schema, then manually validate the results because diagnosis codes, lab values, and clinical notes tell different stories. A "diabetes cohort" defined by ICD-10 codes misses patients documented only in progress notes. A "heart failure readmission risk" query breaks when the admitting diagnosis field is blank but the discharge summary tells the full story.
Clarity's automated cohort identification feature attempts to solve this by combining configurable clinical criteria with natural language processing (NLP) over unstructured text. The engineering decision here isn't just adding NLP—it's exposing the criteria as versioned, testable configuration rather than hardcoded logic.
How the Criteria Engine Works
The cohort builder lets clinical analysts define inclusion and exclusion rules using a structured expression language. Each criterion references standardized data elements—diagnosis codes (ICD-10, SNOMED), medication orders (RxNorm), lab results (LOINC), encounter types—and logical operators (AND, OR, NOT, temporal windows). For example: "Patients with ≥2 HbA1c > 8.0% in the last 12 months AND no endocrinology referral in the last 6 months."
Where it gets practical: the NLP module runs as a parallel pipeline. It ingests clinical notes, discharge summaries, and radiology reports, then extracts entities (conditions, medications, negations, temporality) using a healthcare-tuned model. The extracted assertions are mapped to the same standardized vocabularies the structured criteria use. So a note stating "patient's diabetes is well controlled on metformin" contributes to the diabetes cohort even if no coded diagnosis exists in the structured data.
The key engineering choice: NLP output is treated as evidence, not ground truth. Each extracted assertion carries a confidence score and provenance (document ID, section, character offset). The cohort engine lets you weight structured and unstructured evidence differently per criterion—e.g., require structured lab confirmation for HbA1c but accept NLP-only evidence for "documented patient education."
Integration Reality: FHIR as the Contract, Not the Guarantee
Clarity ingests data via HL7 FHIR R4 bundles from source systems. The platform expects Patient, Encounter, Condition, Observation, MedicationRequest, and DiagnosticReport resources. In practice, the fidelity varies wildly. One Epic instance may populate Observation.valueQuantity with proper UCUM units for every lab; another sends only string values with free-text units. A Cerner feed might include DiagnosticReport.presentedForm (the PDF) but no coded results.
The cohort engine handles this by defining data readiness rules per data element. For each LOINC code you care about, you specify: required FHIR profiles, acceptable value types, unit normalization logic, and fallback to NLP extraction from DiagnosticReport text. This moves the data quality burden from the cohort definition to a reusable catalog—analysts don't rewrite parsing logic for every new cohort.
Required permissions: FHIR read access across the relevant resource types, typically granted via a SMART on FHIR backend service client with system-level scopes. The integration team must map each source system's custom extensions to the platform's canonical data model before cohort evaluation runs.
Worked Example: Identifying Candidates for a Transitional Care Program
Goal: Find patients discharged in the last 30 days with heart failure (HF) who are at high readmission risk and not yet enrolled in the transitional care program.
# Structured criteria (pseudocode)
INCLUDE:
- Encounter: type=inpatient, dischargeDate within [-30d, now]
- Condition: code in (I50.*, I11.0, I13.0, I13.2) # HF ICD-10 cluster
- Observation: code=39156-5 (BMI), value > 35 # obesity comorbidity
- MedicationRequest: code in (loop diuretics, ACEi/ARB, beta-blockers), status=active
EXCLUDE:
- Procedure: code=transitional_care_enrollment
- Encounter: type=emergency, period.start within [-7d, now] # already readmitted
# NLP-augmented criteria
INCLUDE (NLP weight=0.7):
- Entity: "heart failure" OR "HF" OR "systolic dysfunction" OR "diastolic dysfunction"
with negation=false, temporality=current
- Entity: "non-adherent" OR "medication non-compliance" OR "missed doses"
with negation=false
# Risk score threshold (predictive model output)
INCLUDE: risk_score >= 0.65
What happens at runtime: The engine pulls structured data via FHIR, runs the NLP pipeline on all notes for the candidate patients, merges evidence per criterion, applies the weights, and outputs a cohort list with an audit trail showing which criteria each patient met and via what evidence type (structured, NLP, or both). The audit trail is critical for compliance review—you can show exactly why patient X qualified.
Trade-off: Configuration Depth vs. Vendor Dependency
The expression language and NLP weighting are powerful, but they create a specialization trap. Complex cohort definitions become proprietary assets locked into Clarity's engine. If you later migrate to another platform, you're rewriting every cohort from scratch. The mitigation: export cohort definitions as JSON (the platform supports this) and maintain a separate, platform-agnostic specification document that captures the clinical intent, not just the syntax.
Another limitation: NLP model updates. When the vendor upgrades the extraction model, previously extracted assertions don't automatically refresh. You must re-run the NLP pipeline across the corpus, which incurs compute cost and may change cohort membership retroactively. Plan for quarterly re-evaluation windows and version your cohort snapshots.
Actionable Next Steps
- Inventory your cohort definitions today. List every SQL query, registry report, and manual chart review process. Tag each with: clinical owner, refresh frequency, data sources used, and known gaps (e.g., "misses patients documented only in notes").
- Prototype one high-value cohort in Clarity's sandbox. Use the FHIR test harness to validate your source system's data completeness for the required elements before committing to the NLP pipeline.
- Define your NLP confidence thresholds per use case. A clinical trial screen needs higher precision (fewer false positives) than a population health outreach list. Document these thresholds in your data governance policy.
- Schedule a quarterly cohort audit. Compare Clarity's output against a gold-standard manual review of 50–100 patients. Track drift from NLP model updates and source system changes.
The platform's value isn't the NLP alone—it's the discipline it forces: explicit criteria, auditable evidence, and a shared language between clinicians and engineers. Start with the cohort that causes the most manual rework today.
0 replies
A thoughtful contribution can make all the difference. Be the first to share one.