Skip to content
Book Open access

C3Flow: SIMD-Style Concurrent Claude Code Workflow for Scaling Deep Research

Jul 2026 · Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · pp. 5151-5155 · 0 citations · 27 references
Computer Science

TL;DR

C3Flow (Concurrent Claude Code Workflow), a framework that transforms Claude Code from an interactive assistant into a SIMD-style SIMD-style concurrent compute engine, is introduced, demonstrating its effectiveness for production-scale deep research pipelines.

Abstract

Deep research is a retrieval-intensive task that requires iteratively retrieving evidence, reading across sources, and synthesizing source-grounded outputs. In practice, real-world deep research applications are of high workloads that require generating massive reports or conducting large-scale literature surveys. Such applications increasingly require batch processing capabilities, which are missing from traditional chat-oriented agents, limiting throughput when processing large volumes of structurally similar jobs. We introduce C3Flow (Concurrent Claude Code Workflow), a framework that transforms Claude Code from an interactive assistant into a SIMD-style (Single Instruction, Multiple Data) concurrent compute engine. C³Flow treats each agent instance as an isolated, schedulable unit capable of handling declarative multi-step tasks, multi-model routing, and comprehensive trajectory logging. On BrowseComp-zh, C³Flow improves pass@1 from 48.44% to 61.59% and pass@3 from 70.24% to 77.51% compared to standard function calling, while reducing average latency. For multi-hop fact verification, C3Flow achieves a 5.9 speedup over human annotators while maintaining 87.5% accuracy, demonstrating its effectiveness for production-scale deep research pipelines. Code is available at~ https://github.com/RAGenius/C3Flow.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

RAILS: Retrieval-Augmented Incremental LLM Clustering at Scale

Using a Large Language Model (LLM) as the clusterer at production scale is hard: prompts cannot hold the entire label space, and per-document serial processing does not deliver the throughput real workloads require. We present RAILS, a retrieval-augmented incremental LLM clusterer that turns clustering into a simple lo...

Armin Oliya, Aleksandra Sawczuk, Radosław Białobrzeski · 0 citations
#small language model Book Open access Aug 2026

Balancing and Beyond: Communication-Centric Optimizations in Expert Parallelism

EPIC mitigates imbalance via performance-aware expert migration and runtime expert activation, and then improves communication with topology-adaptive transport kernels and fine-grained computation-communication overlap.

Jia-Min Cao, Qingxu Li, Yaozhong Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

vla.simd: Efficient CPU Inference for Language-Conditioned Manipulation

Deploying language-conditioned manipulation without a dedicated GPU requires efficient inference and action chunks that cover the delay between policy queries. We present vla.simd, a CPU inference engine that combines shared SIMD micro-kernels, reusable computation, and target-specific optimization. We relate query lat...

Khanh Duy Nguyen, Hoang M. Truong, A. T. Le · 0 citations
#artificial intelligence Preprint Sep 2026

Efficient Iterative Retrieval with Heterogeneous Batching

Modern information retrieval increasingly employs both embedding and generative models to handle complex queries. However, current serving systems suffer from low throughput and poor GPU utilization because they execute these models in isolation. Coarse-grained partitioning, such as dedicating GPUs to specific tasks, f...

Dohyun Park, Hubertus Franke, Daniel G. Waddington et al. · 0 citations
Preprint Aug 2026

PeakBench: Benchmarking Resource-Aware Tool Invocation in LLM Agents

PeakBench is a benchmark of executable multi-tool workflows with execution-grounded dependency annotations and measured resource profiles that shows that strong logical planning does not reliably translate into safe or efficient execution under resource constraints, and exposes resource information to reduce avoidable...

Zhi-Kai Chen, Xu-Xiang Zhong, Song-Yan Li et al. · 0 citations
#natural language process... Preprint Sep 2026

Beyond Fluent Generation: A CPU Reliability Benchmark for MCP-Style Tool Calling in Sub-2B Small Language Models for Edge Deployment

Resource constrained single-board computers including Raspberry Pi, NVIDIA Jetson Nano, Arduino UNO Q, Orange Pi, and LattePanda motivate on-device small language model (SLM) agents that reduce cloud dependence, improve data locality, and tolerate intermittent connectivity. Model Context Protocol (MCP)-style tool invoc...

Abrar Shahriar Qurat-Ul-Ain Mastoi · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.