Home / Research
BillingBell Sentinel™, in development
Verify AI-generated healthcare billing actions before submission. This page presents the plan for developing and evaluating Sentinel, proposed by BillingBell Healthcare Systems.
Project summary
Artificial intelligence is rapidly being adopted to assign medical codes and prepare insurance claims, yet no independent step routinely verifies AI-generated billing actions before they are submitted. Errors made at scale reach payers as denials and reach patients as incorrect balances, confusing statements and coverage problems that can delay care. Independent and small group practices, which serve many patients in their communities, have the fewest resources to detect these errors.
BillingBell Healthcare Systems proposes to develop and evaluate BillingBell Sentinel™, software that checks each AI-generated claim against the clinical documentation, national coding edits and payer rules before submission, and returns any claim that fails to a human coder with a plain-language explanation. The project will measure whether pre-submission verification reduces billing errors, patient-facing billing harm and administrative burden for small practices.
Relevance to patients and public health
Patients should receive bills that match the care they received. This project will develop a safeguard that catches errors made by AI billing tools before they become denied claims, incorrect patient balances or delays in care, with a focus on the small community practices where many patients receive their care.
Specific aims
Aim 1. Develop and validate Sentinel's error-detection rules and models.
Using de-identified claims and clinical notes, build checks for code-to-documentation match, E/M level, NCCI edits, payer rules, duplicates and patient-balance errors. Validate them against independent review by certified coders, measuring sensitivity and false-hold rate.
Aim 2. Test Sentinel in real-world small-practice workflows.
Deploy Sentinel at pilot practices alongside their existing AI tools. Assess usability, coder review time and fit with daily work, refining the design with input from practice staff and patient advisors.
Aim 3. Evaluate effects on practices and patients.
Compare billing outcomes before and after Sentinel at each pilot site: denial rate, days in accounts receivable, rework time and incorrect patient statements.
Significance
- AI coding tools are spreading faster than methods to verify their output.
- A single systematic AI error can repeat across many claims before it is detected.
- Billing errors become patient burden: unexpected balances, collections and delayed care.
- Small practices bear disproportionate administrative cost and compliance risk.
Innovation
Independent of the AI tool
Sentinel checks output from any AI coding or billing system, rather than being part of it.
Reads the clinical note
It verifies that documentation supports each code, beyond the format checks of traditional claim scrubbers.
Explains every hold
Each flagged claim carries a reason a coder can act on, keeping a person in control.
Patient-balance checks
It looks for errors likely to produce a wrong bill for the patient, not only payer rejections.
Approach
Data and reference standard
De-identified claims and notes from participating practices, under Business Associate and data-use agreements. Certified coders review samples independently to create the reference standard.
Development
Rules based on current CPT®, HCPCS, ICD-10-CM and NCCI guidance, combined with models that compare documentation with assigned codes. Every decision is logged with its reason.
Pilot deployment
Sentinel runs at pilot practices between the AI tool and claim submission. Coders review every held claim and make the final decision.
Evaluation
Pre/post comparison at each site using the measures below, with results reported in full, including missed errors and wrong holds.
Outcome measures
| Measure | Definition | Benefits |
|---|---|---|
| Catch rate | Held claims confirmed wrong by expert coder review | Practices, patients |
| False-hold rate | Held claims that were actually correct | Practices |
| Error mix | Holds by type: code, documentation, modifier, payer rule, duplicate | Practices |
| Denial rate | Denied claims as a share of submitted, before and after | Practices |
| Days in A/R | Average time from submission to payment | Practices |
| Patient-bill errors | Incorrect patient statements prevented | Patients |
| Review time | Minutes of coder review per held claim | Practices |
Protection of patients and data
- Business Associate Agreements with every participating practice
- HIPAA Safe Harbor de-identification before any analysis
- Independent ethics review of study procedures where required
- Data hosted in the United States, encrypted and access-logged
- No change to patient care; a person makes every billing decision
Development phases
Phase 1: Build and validate
Develop core checks and validate them against expert coder review (Aim 1).
Phase 2: Pilot in practices
Deploy at pilot sites, gather staff and patient-advisor feedback, and refine (Aim 2).
Phase 3: Evaluate and share
Measure outcomes and publish findings for practices, payers and researchers (Aim 3).
Sharing results
Findings will be shared openly: summaries for practices and patients on this website, and full results through conference presentations and peer-reviewed publication with research partners.
Partners we are seeking
Pilot practices
Independent and small group practices using, or planning to use, AI in billing.
Research collaborators
Academic and health-system investigators in health services, informatics or patient financial burden.
Patient advisors
Patients and caregivers who can help make bills clearer and fairer.
Funding partners
Organizations supporting safe AI and patient-centered care.
Write to us at research@billingbell.com.
Partner on Sentinel
Join as a pilot practice, research collaborator, patient advisor or funding partner.