Legal AI benchmarks, citation checks, and retrieval-grounding tests primarily evaluate upstream capability: whether a model can answer, extract, or ground a legal task. Deployment asks a different question: whether a particular output remains observable enough to be deployed, reviewed, corrected, or escalated once it enters an organizational workflow. We introduce the Workflow Signal Protocol (WSP), a deployment-layer measurement method for recording workflow observability as a structured workflow-observability record. WSP encodes source status, proposition support, review state, recourse, provenance, and role-scoped disclosure. We validate WSP through controlled stress tests, public legal datasets, documented real-world failures, and a live-output pilot using three general-purpose model application programming interface (API) arms. In the main matched-vocabulary stress test, local formal/substantive routing reduced hidden-risk deployment from 93.5% under calibration-only abstention to 4.0% or below; all 192 pilot outputs were expressible as WSP records. These results support the central claim that deployment-time workflow observability is measurable within the evaluated legal-AI settings; validation in deployed institutional legal workflows remains future work.
Artificial intelligence (AI)-enabled medical device software is increasingly expected to learn, update, explain, retrieve information, draft records, and support clinical reasoning across changing care environments. Regulatory science has responded with lifecycle-oriented tools, including software-as-a-medical-device risk categorization, quality management systems, medical device software lifecycle processes, risk management, post-market monitoring, and predetermined change control plans. These tools are necessary, but they do not by themselves specify who must respond when field experience shows that an AI-supported clinical decision, workflow, validation claim, or planned update no longer fits clinical reality. Documented deployments of sepsis early-warning software in critical care—an external validation that overturned a widely deployed model’s performance claims, the manufacturer’s subsequent model replacement, and a prospective multi-site study linking alert response latency to sepsis mortality—show that such field signals are common, consequential, and unevenly answered. We propose lifecycle answerability as a regulatory science construct for AI-enabled medical device software. It specifies standing, addressee, reason-giving, temporal trigger, and revision pathway. We provisionally define the credible field signal that triggers these obligations, differentiate the construct from established algorithmic accountability frameworks, apply it in parallel to sepsis early warning and large language model documentation, and examine what is distinctive about answerability obligations in critically ill populations. Future work should test answerability through deployment case reviews, post-market signal audits, escalation pathway simulations, and implementation studies. Lifecycle answerability complements existing lifecycle governance by specifying the institutional response architecture through which credible field signals become institutionally actionable across the software lifecycle.