Skip to content
Preprint

Clinical Graph-JEPA: Predictive Patient-State Knowledge Graphs for Cognitive Decision Support

Aug 2026 · 0 citations · 10 references
Computer Science

TL;DR

A clinical knowledge graph construction and refinement framework that combines multi-agent relation proposal, ontology-aware normalization, deterministic evidence scoring, and JEPA-based latent refinement is proposed.

Abstract

Clinical records contain rich evidence about patient state, but converting that evidence into reliable, structured knowledge graphs remains difficult because extraction errors, ontology mismatch, missing relations, and temporal ambiguity can propagate into downstream systems. We propose a clinical knowledge graph construction and refinement framework that combines multi-agent relation proposal, ontology-aware normalization, deterministic evidence scoring, and JEPA-based latent refinement. Rather than treating a clinical knowledge graph as a static extraction artifact, we treat it as a predictive patient-state representation. For each admission, the system constructs an evidence-scored graph from structured MIMIC-IV records and inferred clinical cross-links, then learns to recover held-out clinical relations from the observed graph context. We evaluate the refiner with leakage-free leave-one-out edge recovery (MRR and Hits@k) and held-out batch-mask evaluation (AUC and MRR). To isolate the contribution of discharge-note context, we compare a note-embedding-free configuration with a note-augmented configuration that injects real discharge-note representations only into note-grounded entities. Under the same cohort and evaluation protocol, entity-grounded note injection improves overall leave-one-out MRR by 31% relative improvement.

View source

Similar papers

Aug 2026

CGX: OCR-enhanced knowledge graph retrieval for explainable heart failure analysis.

Initial experiments on heart-failure-focused clinical question answering show that CGX improves evidence retrieval quality and perceived answer reliability over conventional retrieval methods, while reducing total graph construction time by 69.7% under the same input corpus and hardware setting.

Dat Nguyen, Anh N. Le, Binh T. D. Trinh et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Propagating construction-time knowledge quality into medical question answering: A framework grounded in clinical guidelines

Large language models have facilitated knowledge graph (KG) construction from clinical guidelines, but extracted triples vary in structural validity and evidential support. Meanwhile, graph-augmented question answering (QA) systems typically optimize query relevance during retrieval, with limited reuse of quality information produced during KG construction. This creates a disconnect between construction-time quality control and inference-time evidence use. We investigate whether construction-time triple quality can serve as a persistent signal for downstream evidence selection and presentation. We propose a quality-aware framework that models structural conformance (SchemaConf) and evidential support (EvidScore) as complementary dimensions and fuses them into a per-triple quality signal, Q(t). Rather than using quality solely for filtering, the framework retains Q(t) and derived quality tiers as graph attributes and propagates them into quality-weighted subgraph retrieval and tier-conditioned evidence prompting, while preserving passage-level provenance. Experiments on Chinese diabetes clinical guidelines show that the utility of the quality signal is distribution dependent. Under cross-version and cross-model shift, the fused Q(t) provides stronger triple-quality discrimination than either component alone (AUC 0.748 vs. 0.703 for EvidScore and 0.645 for SchemaConf). In guideline-grounded QA, propagating construction-time quality reduces required-knowledge omission from 16.3% to 5.3% and conflicting outputs from 16.3% to 2.7%, with an evidence-grounded precision of 81.6% and near-zero invalid citations. Blinded clinician ratings favor the full framework over no retrieval (4.68 vs. 4.21 on a five-point scale) and approach the oracle condition (4.80), while cross-generator experiments show consistent trends.

Jie Hu, Jun-Jie Wang, Shan Lu et al. · 0 citations
Preprint Aug 2026

Framework for Grounding Healthcare LLMs in a Causal Knowledge Graph: A Cardiovascular Example Pilot

Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.

Ummara Mumtaz, Aimen Noor, Awais Ahmed · 0 citations
Open access Jul 2026

A knowledge graph–driven big data framework for explainable clinical decision support using heterogeneous healthcare data

A scalable, knowledge graph driven big data framework for explainable clinical decision support that unifies heterogeneous healthcare data into a semantically structured representation and provides transparent, traceable decision paths through knowledge graph reasoning addressing key challenges of interpretability and trust in clinical AI systems is proposed.

Ubaid Ul Rehman, Hufsa Mohsin, Ghulam Mustafa et al. · 0 citations
Open access Jul 2026

A Neuro-Symbolic Knowledge Graph and Large Language Model Hybrid Architecture for Multi-Modality Mental Health Counseling

Background: Depression and anxiety are managed largely between clinical visits, yet outpatient care lacks scalable, accountable mechanisms for between-visit support. Large language models converse fluently but fuse clinical reasoning with language generation in one opaque process, so they cannot reliably deliver evidence-based psychotherapy and typically operate outside clinician oversight. Objective: To evaluate C-Mind, a provider-supervised neuro-symbolic system in which a Clinical Knowledge Graph (KG) governs therapeutic decisions for a large language model across eight psychotherapy modalities. Methods: Two simulation regimes addressed eight pre-specified governance questions: a structural validation of KG routing against 117 guideline-anchored vignettes, and a governance battery using progressively disclosing LLM patient agents to evaluate decision traceability, repeatability, provenance auditability, adversarial crisis-detection robustness (277 probes), provider treatment-goal governance, and counselor technique adherence. Crisis detection was additionally validated externally against an independent, clinician-annotated corpus (CRADLE Bench). Results: The KG routed 116/117 vignettes (99.1%) to guideline-appropriate care and detected all 18 high-risk presentations, firing a therapy-suppressing hard halt on 16/18. Adversarial crisis-detection sensitivity was 96.7% and specificity 95.4% (277 probes); on external validation, the system detected 98.5% of 600 dialogues with ongoing suicidal ideation or self-harm at or before the annotator confirming turn. Decisions were 99.1% repeatable, 100% reconstructable per turn, and 100% provenance-auditable across all 354 KG nodes. Provider-set diagnosis, goals, and safety context governed behavior deterministically. Stripped of governance, the same model produced unsolicited clinical monologues on 100% of turns (vs 9% governed) and delivered diagnoses and medication advice the governed system never produced. Conclusions: A neuro-symbolic architecture achieves near-perfect guideline-appropriate routing with a governance profile, traceability, reproducibility, machine-traceable provenance, externally validated crisis detection, and deterministic provider control aligned with requirements for regulated clinical AI.

J. Tao, N. Fenn, H. Parent et al. · 0 citations