Skip to content
View reesemeres-jpg's full-sized avatar
🙃
🩺 RN → Healthcare AI Evaluation
🙃
🩺 RN → Healthcare AI Evaluation

Block or report reesemeres-jpg

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
reesemeres-jpg/README.md

Hi, I’m Meredith Reese, RN 👋

I’m a registered nurse working in healthcare AI evaluation, benchmark design, and trustworthy clinical reasoning.

I design original benchmark suites that test whether AI systems can reason safely through realistic healthcare decisions—not merely produce confident or polished answers.

Healthcare AI Evaluation Portfolio

My work currently focuses on three complementary reasoning domains:

  • operational decision-making under real-world workflow constraints;
  • evidence integrity and eligibility determination; and
  • uncertainty recognition, confidence calibration, and justification boundaries.

Flagship Benchmark Suites

Evaluates whether AI systems can reason through realistic healthcare operations involving competing priorities, staffing limitations, workflow dependencies, clinical risk, scope boundaries, and incomplete information.

Core question:
Can the AI reason through realistic healthcare operations?


Evaluates whether AI conclusions are actually supported by evidence that is sufficient, current, internally consistent, patient-specific, and authoritative.

Core question:
Has the AI earned the right to reach this conclusion?


Evaluates whether AI systems recognize uncertainty, remain within justification boundaries, calibrate confidence appropriately, choose proportionate actions, and define meaningful reassessment.

Core question:
Has the AI exceeded what it can justify?

Evaluation Principles

My benchmark work emphasizes:

  • Evidence over labels or assumptions
  • Object-specific confidence
  • Decision-blocking versus non-blocking uncertainty
  • Safe and proportionate action
  • Clinical and operational realism
  • Reassessment when new evidence emerges
  • Explicit justification boundaries
  • Balanced penalties for both overreach and unnecessary deferral

The safest conclusion is the strongest conclusion the evidence can justify.

Professional Background

My clinical background as a registered nurse informs how I design evaluation tasks involving:

  • Clinical documentation
  • Healthcare workflows
  • Patient safety
  • Care coordination
  • EHR-based reasoning
  • Documentation integrity
  • Operational constraints
  • Multidisciplinary decision-making
  • Human–AI collaboration

Areas of Interest

  • Healthcare AI evaluation
  • LLM evaluation
  • Benchmark and rubric design
  • Clinical reasoning
  • AI safety and reliability
  • Evidence validation
  • Uncertainty management
  • Healthcare operations
  • Clinical documentation
  • Responsible AI development

Connect

Pinned Loading

  1. HORB-Healthcare-AI-Benchmark-Suite HORB-Healthcare-AI-Benchmark-Suite Public

    A healthcare AI evaluation benchmark suite testing operational reasoning, workflow constraints, uncertainty, and executable readiness across realistic clinical operations.

    HTML 2

  2. EIED-Healthcare-AI-Benchmark-Suite EIED-Healthcare-AI-Benchmark-Suite Public

    Evidence Integrity & Eligibility Determination (EIED): A Healthcare AI Benchmark Suite for Evidence Validation, Documentation Integrity, and Eligibility Reasoning.

    1

  3. UCM-Healthcare-AI-Benchmark-Suite UCM-Healthcare-AI-Benchmark-Suite Public

    Uncertainty & Confidence Management (UCM): A healthcare AI benchmark suite for uncertainty recognition, justification boundaries, confidence calibration, proportionate action, and reassessment.

    1