I’m a registered nurse working in healthcare AI evaluation, benchmark design, and trustworthy clinical reasoning.
I design original benchmark suites that test whether AI systems can reason safely through realistic healthcare decisions—not merely produce confident or polished answers.
My work currently focuses on three complementary reasoning domains:
- operational decision-making under real-world workflow constraints;
- evidence integrity and eligibility determination; and
- uncertainty recognition, confidence calibration, and justification boundaries.
Evaluates whether AI systems can reason through realistic healthcare operations involving competing priorities, staffing limitations, workflow dependencies, clinical risk, scope boundaries, and incomplete information.
Core question:
Can the AI reason through realistic healthcare operations?
Evaluates whether AI conclusions are actually supported by evidence that is sufficient, current, internally consistent, patient-specific, and authoritative.
Core question:
Has the AI earned the right to reach this conclusion?
Evaluates whether AI systems recognize uncertainty, remain within justification boundaries, calibrate confidence appropriately, choose proportionate actions, and define meaningful reassessment.
Core question:
Has the AI exceeded what it can justify?
My benchmark work emphasizes:
- Evidence over labels or assumptions
- Object-specific confidence
- Decision-blocking versus non-blocking uncertainty
- Safe and proportionate action
- Clinical and operational realism
- Reassessment when new evidence emerges
- Explicit justification boundaries
- Balanced penalties for both overreach and unnecessary deferral
The safest conclusion is the strongest conclusion the evidence can justify.
My clinical background as a registered nurse informs how I design evaluation tasks involving:
- Clinical documentation
- Healthcare workflows
- Patient safety
- Care coordination
- EHR-based reasoning
- Documentation integrity
- Operational constraints
- Multidisciplinary decision-making
- Human–AI collaboration
- Healthcare AI evaluation
- LLM evaluation
- Benchmark and rubric design
- Clinical reasoning
- AI safety and reliability
- Evidence validation
- Uncertainty management
- Healthcare operations
- Clinical documentation
- Responsible AI development