Current state-of-the-art LLMs as a tool for maritime navigation, which includes both codified rules in the Collision Regulations and uncodified best practices summarized in the concept of ``Good Seamanship'' are explored.
Abstract
Recently, Large Language Models (LLMs) have shown considerable capability for situational understanding, reasoning, and decision making in different domains, most notable in the automotive sector. Therefore, we explore current state-of-the-art LLMs as a tool for maritime navigation, which includes both codified rules in the Collision Regulations (COLREGs) and uncodified best practices summarized in the concept of ``Good Seamanship''. We construct a dataset consisting of 50 diverse, real-world navigation scenarios from AIS data, label scenarios with applicable COLREG rules, recommended actions, and the reasoning for the action. We explore a variety of different LLM architectures and sizes to determine their understanding of maritime navigation tasks as well as evaluate their reasoning capabilities in this domain. The results obtained indicate that the maritime navigation task remains difficult to solve without fine-tuning, even for larger online models.
The correct regulatory interpretation in naval environments is challenging due to the complexity and urgency of decisions based on the International Regulations for Preventing Collisions at Sea (COLREGs). This article presents the development of an intelligent agent named Cognitive Agent for Analysis of Interrelated Problems in Maritime Navigation, hereafter referred to as CAPTAIMN. This agent integrates Large Language Models (LLM) and a Retrieval-Augmented Generation (RAG) architecture to support human decision-making and officer training in safety-critical naval environments. Thus, this work aims to propose a methodology for building an intelligent agent based on LLM and RAG, specifically focused on the assisted and contextualized interpretation of COLREGs. The proposed methodology was evaluated through a quantitative study with 15 maneuvering officers, who assessed 150 responses generated by a local language model using a Likert scale. The results from this phase showed significant approval, with 80% of the responses being rated as 'Agree' or 'Totally Agree' by the officers. These results suggest that the integration of LLM and RAG through CAPTAIMN can provide useful support for both decision-making and tactical training in naval operations.
Gabriel de Sapienza Luna, Arthur Pinheiro de Araújo Costa, Allyson A. da Silva et al.· International Conferences on...· 0 citations
FlightLLM, a prior-guided semantic LLM-based approach for interpretable flight safety analysis that achieves competitive classification performance while generating direct and reasonable explanations for event causes is proposed.
Lang2Graph is presented, an experimental framework for indoor topological graph inference from natural-language navigational instructions that isolates four governing factors: instruction structure, metadata clarity, prompting strategy, and model size and reasoning capability, and Reasoning-aligned open-source models of moderate scale (14B parameters) outperform larger proprietary models on the most challenging instruction categories.
Moamin Ibrahim, Yaqoob Ansari, Khaled A. Harras et al.· International Conference on...· 0 citations
By systematically revealing the strengths and limitations of existing models in aerial-ground collaborative reasoning, AeroGround provides a foundation for developing more capable aerial-ground collaborative embodied intelligence systems.
Shenghong Yi, Lin Zhang, Muzian Li et al.· 0 citations
NavAI is an extensible navigation framework that leverages large language models (LLMs) to support both basic action commands and multi-step goal-oriented navigation through an application-agnostic screenshot-and-control interface and explores optimization strategies for virtual scene understanding and navigation goal decision making.
Jiajie Wang, Sumesh Surendran Letha, M. DiGiovanni et al.· International Conference on...· 0 citations
LAPF is the only evaluated approach that couples every detected hazard to a bounded, metric-neutral corrective action while maintaining near-goal stability, with zero clamp events in both scenarios, whereas CoT prompting increases from 9.7 to 14.0 events.
Yousef Emami, MohammadHossein Homaei, Hao Zhou et al.· 0 citations