<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title>Large language models (LLMs)</title>
    <link>https://www.amazon.science/tag/large-language-models</link>
    <description>Large language models (LLMs)</description>
    <language>en-US</language>
    <lastBuildDate>Fri, 25 Sep 2026 20:36:33 GMT</lastBuildDate>
    <atom:link href="https://www.amazon.science/tag/large-language-models.rss" type="application/rss+xml" rel="self" />
    <item>
      <title>ESPO: Error-structured prompt optimization via diagnose, diversify, and stabilize</title>
      <link>https://www.amazon.science/publications/espo-error-structured-prompt-optimization-via-diagnose-diversify-and-stabilize</link>
      <description>Evolutionary prompt optimizers such as GEPA suffer from prompt bloat: each iteration appends rules and caveats, producing prompts up to 3&amp;#215;longer yet no more accurate. We trace this to three deficiencies - incomplete error observation, limited search diversity, and unreliable selection - and propose ESPO (Error-Structured Prompt Optimization), which decomposes prompt optimization into three phases: Diagnose clusters all training errors into structural patterns in one round; Propose generates candidates via four complementary strategies with independent biases; Select applies bootstrap stability selection. On seven public NLP benchmarks - Tweet, MMLU, GSM8K, HotpotQA, ScoNe, HoVer, and PUPA - ESPO improves average accuracy by +3.76 pp over the state-of-the-art (74.67% vs. 70.91% for GEPA), matching or exceeding GEPA on every dataset while producing prompts 47% shorter (1,004 vs. 1,878 chars) and faster at inference. Cross-model experiments across four additional student models (Gemma 3 12B, Mistral 14B, Qwen3 32B, Claude Haiku 4.5) show ESPO yields the best average accuracy on every model tested, with the largest gap on Qwen3 GSM8K (15.00% &amp;#8594;91.40%). A generalization bound (Appendix) grounds each phase in a corresponding term of the test-time gap, and the ablation confirms a key prediction: adding diversity without bootstrap selection actually hurts performance (&amp;#8722;1.20%).</description>
      <pubDate>Fri, 25 Sep 2026 20:36:33 GMT</pubDate>
      <guid>https://www.amazon.science/publications/espo-error-structured-prompt-optimization-via-diagnose-diversify-and-stabilize</guid>
    </item>
    <item>
      <title>SynthAVE: Scalable synthetic labeling for e-commerce with LLM-arena validation</title>
      <link>https://www.amazon.science/publications/synthave-scalable-synthetic-labeling-for-e-commerce-with-llm-arena-validation</link>
      <description>Fine-tuning large language models (LLMs) for e-commerce attribute extraction requires labeled data representative across thousands of product types, attributes, and multiple languages. This combinatorial scale translates to millions of annotations, rendering human labeling prohibitively costly. While recent work has demonstrated synthetic label generation using LLMs (Negri et al., 2025), deploying such approaches at industrial scale requires integrated quality control mechanisms. We present SynthAVE, a large-scale benchmark for attribute-value verification spanning 12,726 products across 229 product types, 792 attributes, and 4 languages (Spanish, French, Italian, German). To validate synthetic labels at scale, we introduce a multi-LLM arena framework where each sample is evaluated by 21 judge configurations (7 model families x3 prompts), with final labels determined via majority voting; disagreements with the synthetic label are expert-adjudicated and agreement cases are audited on a stratified sample. The majority vote ensemble agrees with expert labels at Cohen&amp;apos;s k= 0.92 (95.0% agreement), while individual judges agree with one another only moderately (Fleiss&amp;apos; k= 0.76)&amp;#8212;by design, since we select for diversity. This demonstrates that diverse models with varying individual judgments aggregate into highly reliable predictions, enabling cost-effective validation at scale while concentrating expert effort on the cases where it changes the label. We estimate the resulting label quality at 97.9%.</description>
      <pubDate>Fri, 25 Sep 2026 20:31:23 GMT</pubDate>
      <guid>https://www.amazon.science/publications/synthave-scalable-synthetic-labeling-for-e-commerce-with-llm-arena-validation</guid>
    </item>
    <item>
      <title>Failing agents leave traces: Telemetry signatures for multi-tier agent evaluation</title>
      <link>https://www.amazon.science/publications/failing-agents-leave-traces-telemetry-signatures-for-multi-tier-agent-evaluation</link>
      <description>Agent evaluation today depends on per-trace LLM-judge inference or human review, too expensive to run on every trace; production systems fall back to sampling a fraction. We find that failing agents leave a detectable behavioral signature in standard observability telemetry: disproportionate effort relative to outcome. We formalize four behavioral failure signatures from agent telemetry, validated on 4,671 traces spanning five agent domains: Compensatory Effort (17&amp;#8211;159% more effort on each domain&amp;#8217;s primary signal, all five domains, p&amp;lt;0.006, the dominant signature everywhere tested); Output Inflation (+38% larger patches on SWE-rebench; direction and reliability vary by domain); Unfocused Orchestration (+5.98&amp;#8211;12.3% higher token entropy, four of five domains); and Correlation Decoupling (effort-signal correlations shift structurally between pass and fail traces, p&amp;lt;0.001, three domains). A simple threshold check over these signals, with no judge, no label, and no model call, already catches 52.9&amp;#8211;61.5% of failures. Across three human-labeled benchmarks, judges miss 6.5&amp;#8211;18.1% of confirmed failures on the two directly comparable ones, and a trained telemetry classifier (AUROC 0.821) recovers 61.5&amp;#8211;75.4% of each judge&amp;#8217;s misses; combining judge and classifier cuts the miss rate from 18.1% to 4.4%, a 4.1&amp;#215; reduction against the weakest of the eight judges tested (GPT-4o-mini); stronger judges start with fewer misses, so the reduction is correspondingly smaller.</description>
      <pubDate>Fri, 25 Sep 2026 19:44:31 GMT</pubDate>
      <guid>https://www.amazon.science/publications/failing-agents-leave-traces-telemetry-signatures-for-multi-tier-agent-evaluation</guid>
    </item>
    <item>
      <title>Characterizing and provisioning the environment bubble in agentic-RL LLM rollout</title>
      <link>https://www.amazon.science/publications/characterizing-and-provisioning-the-environment-bubble-in-agentic-rl-llm-rollout</link>
      <description>Reinforcement-learning (RL) post-training of tool-using LLM agents leaves expensive rollout GPUs idle while each trajectory waits on CPU-side environment work: sandbox cold-start, code execution, retrieval, and external API calls. We term this idle the environment bubble and measure it directly on 8&amp;#215;A100 hardware using a 20 Hz NVML profiler that avoids the kernel-residency pitfall in the utilization counter and attributes each idle instant to one of three states: env-wait, straggler, or pipeline. Our central finding is that the magnitude of the bubble is governed by the environment tool&amp;#8217;s cold-start cost rather than by the RL system. Two real-GPU regimes bracket this behavior. Under an aggressive injected coldstart (warm_s=8 s, representative of heavy tools such as fresh containers, VM or browser spin-up, and a costly env.reset), env-wait is large and decreases from 0.80 at k=4 to 0.23 at k=64. Under a real python:3.11-slim Docker codeexecution backend, whose cold-start we measure at 0.3&amp;#8211;1.2 s, env-wait is small, decreasing from 0.049 to 0.008 over the same range, while effective utilization rises to 0.42&amp;#8211;0.55. The large values therefore require heavy tool cold-start; under realistic sub-second tools only a few percent of idle is recoverable, and the tool latency of the target system should be measured before optimization. Increasing the number of sandbox workers reduces the bubble, but the benefit diminishes as coldstart decreases. We model the tier as a closed finite-source (M/M/k//C) queue; because the model overstates the achievable worker savings, we size k from the measured idle-vs-k curve. A lightweight next-tool-call predictor anticipates calls on real rollout token streams at AUC 0.81 without an additional forward pass. Finally, on real hardware speculative prewarming does not recover meaningful GPU idle in either regime, and a naive simulator (a persistent warm pool evaluated with too few seeds) can indicate otherwise. This is a measurement and reality-check study, and each reported value is labeled by how it was obtained.</description>
      <pubDate>Thu, 24 Sep 2026 21:46:57 GMT</pubDate>
      <guid>https://www.amazon.science/publications/characterizing-and-provisioning-the-environment-bubble-in-agentic-rl-llm-rollout</guid>
    </item>
    <item>
      <title>Validation-gated architecture search for self-organizing multi-agent orchestration</title>
      <link>https://www.amazon.science/publications/validation-gated-architecture-search-for-self-organizing-multi-agent-orchestration</link>
      <description>Multi-agent architectures today are hand-designed, generated one-shot, or evolved through unconstrained code search&amp;#8212;none of which treats the number of agents as a searchable parameter, and none of which can discover that a single agent is optimal. We argue agent count and coordination topology should be optimized with the same discipline applied to neural architectures: bounded edits, held-out validation gating, and bidirectional search that can both add and remove agents without information loss. OrchOpt is a validation-gated architecture search over a unified design space where single-agent and multi-agent configurations compete directly &amp;#8212; with bidirectional traversal enabled by knowledge preserving merge. A proposer model turns failure analysis into bounded merge/split/skill edits on the current architecture, and an edit is accepted only when it strictly improves validation. Knowledge preserving merge&amp;#8212;where absorbed agents&amp;#8217; expertise becomes activatable skills&amp;#8212;makes the search fully reversible at zero additional inference cost. Across nine benchmarks spanning math, knowledge QA, code generation, reading comprehension, multi-hop reasoning, negotiation, and research: the optimizer, initialized with the same baseline design, simplifies to 1 agent on 5 domains and elaborates to 2&amp;#8211;3 agents on 4 domains. It finds architectures scoring +20pp over ADAS on MGSM, +16.7pp over AFlow on HotpotQA, and +27.5pp over MARBLE&amp;#8217;s published-best topology on Bargaining&amp;#8212;converging in 7&amp;#8211;42 candidate evaluations.</description>
      <pubDate>Thu, 24 Sep 2026 21:39:52 GMT</pubDate>
      <guid>https://www.amazon.science/publications/validation-gated-architecture-search-for-self-organizing-multi-agent-orchestration</guid>
    </item>
    <item>
      <title>Chart-RL: Policy optimization reinforcement learning for enhanced visual reasoning in chart question answering with Vision Language Models</title>
      <link>https://www.amazon.science/publications/chart-rl-policy-optimization-reinforcement-learning-for-enhanced-visual-reasoning-in-chart-question-answering-with-vision-language-models</link>
      <description>The recent advancements in Vision Language Models (VLMs) have demonstrated progress toward true intelligence requiring robust reasoning capabilities. Beyond pattern recognition, linguistic reasoning must integrate with visual comprehension, particularly for Chart Question Answering (CQA) tasks involving complex data visualizations. Current VLMs face significant limitations in CQA, including imprecise numerical extraction, difficulty interpreting implicit visual relationships, and inadequate attention mechanisms for capturing spatial relationships in charts. In this work, we address these challenges by presenting Chart-RL, a novel reinforcement learning framework that enhances VLMs&amp;apos; chart understanding through feedback-driven policy optimization of visual perception and logical inference. Our key innovation includes a comprehensive framework integrating Reinforcement Learning (RL) from Policy Optimization techniques along with adaptive reward functions, that demonstrates superior performance compared to baseline foundation models and competitive results against larger state-of-the-art architectures. We also integrated Parameter-Efficient Fine-Tuning through Low-Rank Adaptation (LoRA) in the RL framework that only requires single GPU configurations while preserving performance integrity. We conducted extensive benchmarking across open-source, proprietary, and state-of-the-art closed-source models utilizing the ChartQAPro dataset. The RL fine-tuned Qwen3-VL-4B-Instruct model achieved an answer accuracy of 0.634, surpassing the 0.580 accuracy of the Qwen3-VL-8B-Instruct foundation model despite utilizing half the parameter count, while simultaneously reducing inference latency from 31 seconds to 9 seconds. We also performed comprehensive comparative analysis of Chain of Thought (CoT) reasoning performance across those models, demonstrating that the Chart-RL framework significantly enhanced the visual reasoning.</description>
      <pubDate>Thu, 24 Sep 2026 21:00:41 GMT</pubDate>
      <guid>https://www.amazon.science/publications/chart-rl-policy-optimization-reinforcement-learning-for-enhanced-visual-reasoning-in-chart-question-answering-with-vision-language-models</guid>
    </item>
    <item>
      <title>From trace entropy to coordination control: Diagnosing and intervening in multi-Agent LLM systems</title>
      <link>https://www.amazon.science/publications/from-trace-entropy-to-coordination-control-diagnosing-and-intervening-in-multi-agent-llm-systems</link>
      <description>Coordination interventions for multi-agent LLM systems are regime-dependent: identical guidance yields opposite effects at different decoding temperatures, improving accuracy in the deterministic regime but degrading it in the stochastic regime. The reversal traces to over-regularization: exploration-promoting interventions add explicit entropy to policies already receiving implicit entropy from stochastic decoding. The observation motivates a decomposition of the intervention space into regime-dependent types (exploration, commitment) and regime-invariant types (consistency, disruption). We present TRAC (Trace-based Regime-Aware Coordination), a system that implements the regime-appropriate composition through two complementary mechanisms: static self-trend constraints for failures the agent can self-diagnose, and a dynamic callback monitor for failures invisible to the agent. Ablation shows that neither component alone reliably outperforms the uncontrolled baseline, but TRAC addresses their disjoint failure modes jointly. On the GAIA benchmark across two stochastic seeds, TRAC achieves 82.0% versus 75.8% vanilla (+6.2%, McNemar p=0.004).</description>
      <pubDate>Thu, 24 Sep 2026 17:55:56 GMT</pubDate>
      <guid>https://www.amazon.science/publications/from-trace-entropy-to-coordination-control-diagnosing-and-intervening-in-multi-agent-llm-systems</guid>
    </item>
    <item>
      <title>PersonaJudge: Simulating individual human preference judgments with evaluator-specific demonstration data</title>
      <link>https://www.amazon.science/publications/personajudge-simulating-individual-human-preference-judgments-with-evaluator-specific-demonstration-data</link>
      <description>Large language models increasingly serve as judges in AI evaluation, but current approaches rely on consensus preferences that ignore individual evaluator variation. We propose a novel simulation approach that combines categorical judgments with evaluator-specific auxiliary data&amp;#8212;retrospective reasoning traces and interface telemetry&amp;#8212;to enable LLM-based simulation of individual evaluators via in-context learning. We conduct a systematic empirical study of this approach using multi-facet data from 32 trained annotators across 4,200 preference judgments in a 4&amp;#215;4&amp;#215;4 factorial design. Our key findings: (1) The simulation approach achieves up to 9.9 percentage point improvements over the Base Judge; (2) Reasoning traces provide the largest gains with higher collection efforts, while interface telemetry often hurts rather than helps performance despite being cheaper to collect. (3) Simulation difficulty is systematic, predicted by an evaluator&amp;apos;s neutral usage (most clearly on Helpfulness) and divergence from consensus; the neutral-usage tendency&amp;#8212;rather than simulatability itself&amp;#8212;is the cross-task-stable property (r= 0.728 ). These results establish both the potential and limits of evaluator-specific auxiliary data for personalized evaluation, offering methodological insights for scaling individual-aware AI assessment.</description>
      <pubDate>Mon, 21 Sep 2026 17:54:03 GMT</pubDate>
      <guid>https://www.amazon.science/publications/personajudge-simulating-individual-human-preference-judgments-with-evaluator-specific-demonstration-data</guid>
    </item>
    <item>
      <title>Deploying programmatic tool calling with pre-execution validation for production agentic systems</title>
      <link>https://www.amazon.science/publications/deploying-programmatic-tool-calling-with-pre-execution-validation-for-production-agentic-systems</link>
      <description>Production tool calling for agentic systems must satisfy four operational requirements: lower per-invocation cost, low latency, adaptability to evolving tool catalogs, and the ability to catch errors before execution. The ReAct paradigm keeps the large language model (LLM) in the loop, but at production scale, reprocessing intermediate results inflates cost and latency and degrades answer quality. Programmatic tool calling (PTC) offers an alternative; the LLM generates a single script that resolves dependencies and extracts relevant fields in one inference pass. PTC is a harder generation task, however, because the LLM must produce code that handles complex, variable responses without observing intermediate results, and existing remedies each fail at least one of the four requirements. We describe a production deployment of PTC serving 26,528 invocations, paired with a three-layer abstract syntax tree (AST)-based validator that rejects invalid scripts before execution and requires no training data, preserving adaptability to evolving tool catalogs. This reduced average input tokens by 77%, latency by 21%, and increased the answered-query rate by 5.6 percentage points across 19 production agents.</description>
      <pubDate>Mon, 21 Sep 2026 17:43:27 GMT</pubDate>
      <guid>https://www.amazon.science/publications/deploying-programmatic-tool-calling-with-pre-execution-validation-for-production-agentic-systems</guid>
    </item>
    <item>
      <title>TRACE : Traceable root-cause analysis with calibrated evidence-grounded agents</title>
      <link>https://www.amazon.science/publications/trace-traceable-root-cause-analysis-with-calibrated-evidence-grounded-agents</link>
      <description>Root Cause Analysis (RCA) for defects in large e-commerce enterprises is manual and slow: a single defect spanning thousands of microservices and SOPs takes experts one to two weeks to diagnose. Generic agentic RCA frameworks fail on this workload because they stop at proximate causes, cannot verify numeric claims, ignore past reviewer decisions, and return enumerated hypotheses rather than the quantified narratives business teams act on. We present TRACE , an end-to-end agentic RCA system that combines supervised hypothesis-driven investigation with recursive why-drilling, an evidence-traceable multi-tool suite with stable artifact identifiers, a feedback retrieval and ingestion pipeline, and a two-layer confidence framework combining six mechanistic checks with an LLM judge - including a novel numerical-traceability score. Across several thousand cohorts spanning eight use cases in three metric categories (cost, quality, and inventory), TRACE achieves 45.5% expert acceptance; on a 240-cohort head-to-head benchmark, 47.1% vs. 28.8% for a matched ReAct baseline ( +18.3pp, p&amp;lt;0.001) with calibrated confidence (AUC 0.819, ECE 0.041).</description>
      <pubDate>Mon, 21 Sep 2026 16:05:08 GMT</pubDate>
      <guid>https://www.amazon.science/publications/trace-traceable-root-cause-analysis-with-calibrated-evidence-grounded-agents</guid>
    </item>
  </channel>
</rss>
