Validation and beta status
Clinical Corvus is in limited beta. Current validation primarily covers software contracts, traceability, policies, and internal evaluations. It does not demonstrate prospective clinical effectiveness, comparative superiority, or regulatory authorization.
What has been verified
| Area | Available evidence |
|---|---|
| Runtime identity | The health check can expose version and commit to detect a stale service. |
| Canonical pipeline | Governed tasks use /api/agents/tasks/* and produce the shared response contract. |
| Response rails | Canonical responses can carry ClinicalEvidencePacket, PatientContextManifest, and CorvusProvenance. |
| Evidence | Internal evaluations inspect citations, admitted or rejected evidence, and claim-to-source binding. |
| Tenant | Migrations and services apply tenant scoping; covered PostgreSQL tables receive RLS policies. |
| External sharing | Third-party PHI sharing is disabled by default, and public queries receive specific checks. |
These items confirm mechanisms and contracts. They do not confirm uniform coverage across every application path.
Open gates
| Gate | What remains to be shown |
|---|---|
| Clinical utility | Complete responses need systematic specialist review. |
| Source support | Claim coverage must remain consistent across questions and execution paths. |
| Longitudinal handoff | Continuity between shifts and episodes needs operational validation. |
| Product surfaces | Monitor, Plan, Workspace, and Clipboard must preserve context, uncertainty, and traceability end to end. |
| Operations | Latency, availability, support, and incident response must be measured in a pilot. |
| Clinical effectiveness | No prospective improvement in outcomes or care time has been demonstrated. |
Interpreting a response
The interface may show states equivalent to ready for review, partial, insufficient evidence, blocked, or mandatory review. These states guide the user’s next step. They do not validate the clinical content.
Next validation stage
The next stage should combine complete clinical review packets, replay of the screens clinicians use, controlled pilots with predefined measures, and an operational audit of tenant controls, data sharing, billing, logs, and incidents.
Results and dates should be published when a stable, reviewable package exists, not when only an isolated benchmark is available.