Skip to content
Review

Generative AI in the Care of Older Adults: A Position Statement From the American Geriatrics Society.

Jul 2026 · Journal of The American Geriatrics Society · 1 citation · 28 references
Medicine

TL;DR

GenAI should augment, not replace, clinical judgment and relational care, and responsible use in geriatrics requires transparency, clinician-in-the-loop oversight, validation in older adult populations using age-relevant outcomes, and governance safeguards to protect dignity, safety, and equity.

Abstract

Background

Generative artificial intelligence (GenAI), particularly large language models (LLMs), is being integrated into healthcare documentation, decision support, patient education, administrative workflows, and emerging agentic systems capable of initiating clinical and operational actions. While GenAI may reduce clinician burden and support person-centered care, it also introduces risks such as misinformation, algorithmic bias, privacy harms, errors of omission, and automation over-reliance. These risks may be amplified for older adults because core geriatrics care tasks, such as goals-of-care discussions, capacity-sensitive consent, polypharmacy and deprescribing, and functional and cognitive assessment in the setting of multimorbidity, may be underrepresented in training data and are high-stakes in practice.

Methods

The American Geriatrics Society (AGS) convened an interdisciplinary working and writing group, reviewed relevant literature and reports, incorporated input from multiple AGS committees, and completed review and approval through AGS committee processes in March 2026.

Results

This AGS position statement translates core geriatrics principles, person-centered care, equity, shared decision-making, and promoting independence into actionable recommendations for clinicians, health systems, developers, policymakers, older adults, and care partners. Recommendations address ethical use, clinical integration and oversight, transparency and documentation, privacy protections, governance across the AI lifecycle, including pre, during, and post-deployment monitoring and accountability, and priority geriatrics use cases.

Conclusion

GenAI should augment, not replace, clinical judgment and relational care. Responsible use in geriatrics requires transparency, clinician-in-the-loop oversight, validation in older adult populations using age-relevant outcomes, and governance safeguards to protect dignity, safety, and equity.

View source

Similar papers

Review Open access Aug 2026

Safety in the Age of Artificial Intelligence: Evaluating Large Language Model Adherence to Antithrombotic Medication and Regional Anesthesia Guidelines

The release of the 2025 American Society of Regional Anesthesia (ASRA) 5th edition guidelines for regional and neuraxial anesthesia procedures for patients receiving antithrombotic medications introduced complex, patient-specific hold times and resumption protocols. As clinicians increasingly utilize large language models (LLMs) as clinical decision support tools, the reliability of these models remains largely unvalidated. This study evaluates the accuracy of two of the foremost LLMs, ChatGPT (OpenAI, San Francisco, CA) and Google Gemini (Google DeepMind, London, UK), in adhering to these new gold-standard safety guidelines. Twenty-five standardized clinical vignettes were developed. Each vignette featured a patient on a specific anticoagulant (e.g., rivaroxaban, apixaban, dabigatran) requiring a neuraxial or regional anesthetic procedure (stratified by high-risk vs. low-risk). Variables included renal function, dose frequency, and procedural urgency. Prompts were submitted to the latest publicly available ChatGPT and Google Gemini models with separate instructions to provide hold and resumption times. LLMs were queried to ensure familiarity with 2025 ASRA guidelines prior to submission of prompts. Responses were graded against the 2025 ASRA guidelines by independent reviewers. Response adherence was categorized as: 1. concordant (100% match); 2. conservative error (LLM recommended time was longer than required); 3. dangerous error (recommended time was shorter than required, a critical safety violation); or 4. omission (no specific timeframe provided). ChatGPT achieved a concordance rate of 64%, compared to Gemini at 62%. However, the models displayed distinct error profiles. ChatGPT produced "dangerous errors" in 20% of evaluations and failed to specify a time in 16% of cases. In contrast, Gemini’s dangerous error rate was lower at 12%, but it demonstrated a significant "conservative error" rate of 22%, compared to 0% for ChatGPT. A chi-square test indicated that the difference in dangerous error rates between the two models was not statistically significant. Regardless, notable differences in error character and response completeness were observed. Gemini was more consistent in providing specific timeframes in 96% of prompts compared to 84% for ChatGPT. Variability of responses between users was determined via a chi-square test of homogeneity and was found to be statistically significant only for Gemini. While both models demonstrated moderate guideline awareness, their failure modes differed meaningfully. ChatGPT's errors were predominantly dangerous underestimations and omissions, while Gemini exhibited a conservative bias, overestimating hold times when a match was not achieved. Significant limitations exist regarding generalizability of these results. Only two major LLM models were tested. Other, more clinically oriented models exist, and newer versions of both ChatGPT and Gemini are consistently being released. These may improve LLM adherence to clinical guidelines. Despite these limitations, this study highlights the dangers of utilizing LLMs for periprocedural anticoagulation decision-making in regional and neuraxial anesthesia. Both models produced potentially dangerous recommendations in 12-20% of scenarios. Ongoing evaluation of LLM adherence to evolving guidelines remains essential, and clinicians must exercise extreme caution when utilizing these tools in patient care.

Conner M. Willson, Birpartap S. Thind, Jay Srinivas et al. · 0 citations
Review Open access Aug 2026

Beyond Algorithms: Human-Centered Explainability for Clinicians, Patients, and Caregivers

Two short vignettes and a design guide to help human factors researchers create explanation systems that are practical and trustworthy are presented to show how explainability can become part of the care system instead of being treated as a separate technical feature.

T. Mamun, Laurie Novak, M. Salwei · 0 citations
Open access Aug 2026

Artificial Intelligence and Health Care for Older Adults, a Future for All of Us.

We are racing towards a time wherein the largest demographic of the population will be comprised of older adults and where it is now also possible to measure, monitor, and diagnose long term risk factors and provide effective and valuable preventative interventions. Thanks to the achievements of modern health care, we have reached a tipping point in the need for a robust geriatric workforce. Through the implementation of thoughtful, deliberate, and collaborative processes, Artificial Intelligence can play a significant role in preparing individuals in both mind and body for healthy aging. On May 1, 2025, the Johns Hopkins Artificial Intelligence and Technology Collaboratory (JH AITC) hosted a one-day summit in Washington, DC, to explore the future and potential uses of Artificial Intelligence (AI) for older adults. Presenters and participants discussed the challenges facing the implementation of this technology into the U.S. health care system and the opportunities recent advances are affording in improving the quality of life for older adults.

Peter M. Abadir, Phillip H. Phan, M. Aliberti et al. · 0 citations
Review Open access Aug 2026

Large Language Models and Medical AI Systems for Healthcare Diagnosis: A Systematic Review

Although promising, LLM-based systems are not yet reliable enough for autonomous medical diagnosis, and multiple recommendations for future research are contained to ensure a high level of safety, transparency, and clinical applicability for LLMs and other AI/ML-related technologies and devices.

M. U. K. Gunawardhna, Pirunthavi Wijikumar, D. Weerasinghe · 0 citations
Open access Aug 2026

From Algorithm to Bedside: A Clinician's Framework for AI in Practice

Abstract Artificial intelligence (AI) is entering clinical practice faster than the evidence base supporting it. Clinicians, who remain the licensed decision-makers at the bedside, increasingly find themselves as end-users of tools whose strengths, failure modes, and external validity they have had no opportunity to assess. This article offers a practical framework that does not require coding or mathematical literacy. We outline how AI is built, validated, deployed, and monitored, and where each phase typically goes wrong. We propose seven questions that clinicians can run through to evaluate any clinical AI tool in the time it takes to read an abstract, alongside a traffic-light schema for matching oversight to risk and a short list of demands clinicians should make of vendors and institutions. We then examine the deeper questions of equity, accountability, and the therapeutic relationship that AI is now forcing into view. AI literacy belongs alongside biostatistics and evidence-based medicine as a core clinical competency.

Alaa Abdelqader, M. Alkhateeb, Abdullah Al-Marrawi et al. · 0 citations
Review Jul 2026

Can you trust the algorithm? A clinical exploration of AI models in perioperative care

Artificial intelligence (AI) is increasingly influencing perioperative care by enhancing decision support, resource management, documentation, image analysis and patient safety. This article presents a narrative conceptual review of four commonly encountered AI model types: random forests, support vector machines, convolutional neural networks and large language models. It explores their clinical applications, interpretability, limitations and relevance to perioperative and surgical settings. Particular emphasis is placed on model explainability, evaluation metrics such as accuracy, F1 score and area under the receiver operating characteristic curve, hallucination risk and the challenges associated with opaque ‘black box’ systems, including non-interpretable algorithmic frameworks. The discussion is contextualised within the UK’s regulatory and governance frameworks, including the National Health Service Digital clinical safety standards, Medicine and Healthcare products Regulatory Agency requirements, National Institute for Health and Care Excellence evaluation tiers and data protection under UK General Data Protection Regulation. By improving clinician understanding of how AI systems function, fail and should be governed, this article aims to support informed, safe and ethical adoption of AI technologies in perioperative environments.

Kevin Patrick Cairns · 0 citations