Which API for general real-time LLM agents [D]
TL;DR what is the right interface for writing custom real-time LLM agents? I ve noticed many projects (academic, startups, pet, etc) are trying to make LLMs work in real-time world where things happen
At a glance
- reddit.com: Which API for general real-time LLM agents [D]
- arxiv.org: Evaluating Test-Time Scaling of General LLM Agents
- arxiv.org: SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time
The story
reddit.com: TL;DR what is the right interface for writing custom real-time LLM agents? I ve noticed many projects (academic, startups, pet, etc) are trying to make LLMs work in real-time world where things happen while the model is thinking. And the agents need to be fast, but also have some way to new stimuli while they think, like when a user interrupts a voice agent, or browser pop-ups, etc. And these things already exist in principle, e.g. most robotics agents, real-time assistants [videoLLM] [gpt4o] , video calls [wan-streamer] , also [mid-turn steering] in Codex where you can add clarifications while the agent works, etc. So it looks like LLMs can already do real-time interaction for specific scenarios --- I was wondering if it can be generalized so developers can write their own real-time agents as we write custom MCP tools now . So, we would need some sort of reusable building blocks that can be combined into, say, a coding agent that helps you debug in real time, or a deep research agent that talks to you and reacts to your feedback mid-search, or a voice-controlled calendar/spreadsheet helper, etc. A bit silly, my point is: there are applications not covered by current models. I was wondering: if someone were to build this, what interface would you use for building async agents in general? Vendors technically have a real-time APIs [ 1 , 2 ] , but those are for coding. The closest thing I found about general async agents is [AsyncLLM preprint 2609.35427] (disclaimer: I know the authors), where the programmer uses asyncio to write coroutines with shared memory blocks, which is how agents communicate (image below from that preprint). img from https://huggingface.co/papers/2609.35427 However, that still assumes the programmer writes a low-level inference pipeline, so perhaps there s a more natural way that would let openai/anthropic/etc expose some API that would let developers run general async agents for their use case and feed them with real-time streams of events. [ba
arxiv.org: arXiv:2602.18998v2 Announce Type: replace Abstract: LLM agents are increasingly expected to operate as general-purpose systems that resolve real-world user requests, yet their dynamic scaling behavior in realistic environments remains poorly understood. In this paper, we systematically investigate two principal test-time scaling axes of LLM agents: sequential scaling through extended interaction and parallel scaling through trajectory sampling. We first introduce a realistic benchmark that provides one unified framework for evaluating LLM agents across search, coding, reasoning, and tool-use domains, more faithfully reflecting the heterogeneity of real-world deployments. Evaluating ten leading LLM agents reveals substantial performance degradation when transitioning from domain-specific evaluations to this realistic setting. Building on this foundation, we progressively scale test-time compute along fine-grained increments to characterize the performance upper bound. We find that neither scaling axis can consistently yield meaningful gains from additional test-time compute in realistic environments, a phenomenon we attribute to two fundamental limitations: the scaling plateau that bottlenecks sequential scaling and the verification gap that undermines parallel scaling. Code is publicly available at https://github.com/cxcscmu/General-AgentBench.
arxiv.org: arXiv:2609.36580v1 Announce Type: new Abstract: LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCoEvo jointly improves two complementary safety capabilities at different timescales: S-Harness rapidly externalizes recent runtime experience into updatable explicit safety knowledge that can promptly influence subsequent tasks, while GuardVPO internalizes accumulated runtime safety experience over a longer timescale into parametric risk-judgment capabilities. By combining short-term rapid adaptation with long-term capability consolidation, SafeCoEvo continually improves the agent's safety capabilities, reducing the unsafe outcome rate by 10.05% while improving the task success rate by 12.15% over the strongest baseline, thereby achieving simultaneous gains in safety and task utility.
arxiv.org: arXiv:2609.35911v1 Announce Type: cross Abstract: Adapting large language model agents during deployment requires not only retaining past experience, but also turning new observations into timely guidance. Many test-time learning methods, however, acquire knowledge from completed episodes. Feedback from an ongoing interaction may therefore not be distilled into knowledge soon enough to help the next decision. Acquiring knowledge at the granularity of individual transitions could reduce this delay, but raises a separate challenge: a rule that is useful within one episode may not be reliable enough to guide future episodes. Waiting for validation can forfeit immediate benefits, whereas unrestricted reuse can propagate accidental or misattributed guidance. We introduce StepLearn, a nonparametric framework that separates immediate use from persistent trust. It turns informative transitions into hypotheses that can guide the next step, while requiring prospective validation before reuse across episodes. Their predicted effects are checked against subsequent observations outside the source episodes, and only sufficiently supported hypotheses become available for persistent guidance. This process updates external knowledge while keeping all model parameters fixed. Over five rounds on WebArena-Lite and ALFWorld, StepLearn achieves average success rates of 59.9% and 84.0% with GPT-5-mini, and 57.8% and 88.1% with Qwen3.5-35B-A3B, respectively. It outperforms EvoTest, the strongest baseline, by 2.2-12.7 percentage points across the four settings. Learning dynamics further shows that these gains are not restricted to the final repetition, with advantages already present on first task attempts in most settings.