<?xml version="1.0" encoding="UTF-8"?>
<rss xmlns:atom="http://www.w3.org/2005/Atom" version="2.0">
  <channel>
    <title>Knowledge graphs</title>
    <link>https://www.amazon.science/tag/knowledge-graphs</link>
    <description>Knowledge graphs</description>
    <language>en-US</language>
    <lastBuildDate>Tue, 08 Sep 2026 15:50:07 GMT</lastBuildDate>
    <atom:link href="https://www.amazon.science/tag/knowledge-graphs.rss" type="application/rss+xml" rel="self" />
    <item>
      <title>SEGRA: A structured experience guided reasoning agent for property graph question answering</title>
      <link>https://www.amazon.science/publications/segra-a-structured-experience-guided-reasoning-agent-for-property-graph-question-answering</link>
      <description>Enterprise IT support knowledge graphs capture rich relationships among cases, users, devices, symptoms, taxonomic categories, root causes, and historical resolutions. Yet querying them in Gremlin requires knowledge of graph schemas, traversal semantics, edge directionality, and property-graph-specific constraints, making them difficult for non-expert operators to use. We introduce SEGRA, an experience-guided agent for enterprise text-to-Gremlin question answering. SEGRA integrates intent routing, schema- and taxonomy-grounded query generation, multi-shot decomposition, execution-aware verification, and a curriculum-bootstrapped skill library that reuses verified query patterns. On an enterprise IT support benchmark, SEGRA achieves a 7.0 &amp;#215;higher mean judge score than backbone-only chain-of-thought prompting. Its skill library further reduces LLM calls by 20% and dollar cost by 18% relative to SEGRA without skills, while preserving answer quality. These results show that schema-grounded agent design and reusable execution experience improve both accuracy and efficiency for enterprise graph QA.</description>
      <pubDate>Tue, 08 Sep 2026 15:50:07 GMT</pubDate>
      <guid>https://www.amazon.science/publications/segra-a-structured-experience-guided-reasoning-agent-for-property-graph-question-answering</guid>
    </item>
    <item>
      <title>Pairwise ranking outperforms single-action RL for offline explanation selection: A practical lesson</title>
      <link>https://www.amazon.science/publications/pairwise-ranking-outperforms-single-action-rl-for-offline-explanation-selection-a-practical-lesson</link>
      <description>We report a practical lesson from building a GPU-free explainable-recommendation serving stack: explanations are pre-generated offline into a per-item candidate pool, and a small CPU-resident model selects one at request time. Every candidate in the pool of size K carries an offline BERTScore-F1 label against a reference explanation, so we compare a pairwise learning-to-rank model (LightGBM LambdaRank) against a single-action RL formulation (DPO) and a distilled selector on the XRec Google Local benchmark (2,958 pairs, 5 seeds). LambdaRank outperforms both by 0.019&amp;#8211;0.025 F1, a gap more than fifteen times the across-seed standard deviation. The advantage appears structural rather than algorithmic: LambdaRank&amp;apos;s objective consumes the label of every candidate in the pool, whereas DPO&amp;apos;s rollout only observes the label of the sampled side of each preference pair. A separate candidate-source design, selecting from knowledge-graph-grounded paths instead of the cached pool, trades this reference alignment for a Unique-Sentence Ratio (USR) of 1.000. We also encountered two failure modes: further RL fine-tuning on top of an already-distilled policy regresses F1, and an end-to-end RL fine-tune of the generator itself reward-hacks the metric within a few hundred steps. The practical takeaway is that whenever an offline metric can label every candidate in a fixed action set, a pairwise learning-to-rank baseline is worth evaluating before reaching for single-action RL. One caveat: LambdaRank&amp;apos;s training signal differs in form from the evaluation metric we report, since it trains on quintile-binned labels while DPO trains on a blended reward, though both derive from the same underlying BERTScore-F1; a fully decorrelated evaluation remains future work.</description>
      <pubDate>Tue, 18 Aug 2026 16:23:44 GMT</pubDate>
      <guid>https://www.amazon.science/publications/pairwise-ranking-outperforms-single-action-rl-for-offline-explanation-selection-a-practical-lesson</guid>
    </item>
    <item>
      <title>Is GraphRAG needed? From basic RAG to graph-/agentic solutions with context optimization</title>
      <link>https://www.amazon.science/publications/is-graphrag-needed-from-basic-rag-to-graph-agentic-solutions-with-context-optimization</link>
      <description>As advanced RAG variants like GraphRAG and Agentic RAG emerge, one leading question is when and how to use them. Here, we introduce a framework for different RAG scenarios evaluation and comparison on semi-structured knowledge bases, including regular RAG, GraphRAG, Modular RAG and Agentic RAG. We provide implementation for 9 standardized RAG scenarios, and conduct experiments for a comprehensive comparison. These scenarios are designed for real use cases regarding data and domain restrictions, spanning from simple document-based retrieval to advanced features such as hybrid text-graph retrieval, integration with computed or pre-defined domain knowledge graphs, agentic multi-step planning, and agent-graph integration. Besides, we present a novel context engineering method for GraphRAG and Agentic RAG, addressing the context/memory overflow issues, efficiently managing text and graph retrievals with new representations and agentic loop design, leading to 19%-53% reduction on token usage. Moreover, further analysis identifies a retrieval-generation gap where expanded retrieval does not proportionally improve generation quality, suggesting retrieval-oriented metrics overstate advanced retrieval benefits. This work provides data-driven insights on when and how to use them for building production-ready intelligent RAG systems.</description>
      <pubDate>Fri, 24 Jul 2026 15:36:41 GMT</pubDate>
      <guid>https://www.amazon.science/publications/is-graphrag-needed-from-basic-rag-to-graph-agentic-solutions-with-context-optimization</guid>
    </item>
    <item>
      <title>Diagnostic knowledge graphs: Automated benchmark construction and deterministic evaluation for multi-step reasoning agents</title>
      <link>https://www.amazon.science/publications/diagnostic-knowledge-graphs-automated-benchmark-construction-and-deterministic-evaluation-for-multi-step-reasoning-agents</link>
      <description>Evaluating multi-step diagnostic reasoning in LLM agents remains an open problem. When cause labels are extracted from resolved operational cases (customer-service tickets, incident reports, clinical notes), the resulting gold standards exhibit extreme vocabulary explosion&amp;#8212;5,076 unique cause strings from 2,196 tickets on a single symptom, 92% appearing only once&amp;#8212;making LLM-as-judge protocols variance-prone (&amp;#177;2&amp;#8211;3pp inter-run) and longitudinal monitoring impossible. We argue that building reproducible diagnostic benchmarks and building effective diagnostic agents are dual problems solved by the same artifact&amp;#8212;a canonical knowledge structure that normalizes evaluation gold-standards and constrains agent hypothesis spaces simultaneously. We instantiate this duality as Diagnostic Knowledge Graphs (DKGs): hierarchical cause trees with frequency priors, built automatically from resolved tickets via LLM extraction, two-pass BERTopic clustering, and LLM merge, with optional per-node SQL grounding against operational data. The pipeline compresses 5,076 cause strings to 73 canonical clusters (97% coverage) and scales to 284 symptom families from 14,953 tickets without manual curation. A domain-expert validation shows the resulting normalization achieves 92% accuracy&amp;#8212;exceeding the LLM judge&amp;apos;s 80% agreement with the same expert&amp;#8212;while enabling deterministic scoring (&amp;#963;=0 inter-run variance). Under a fair semantic-judge protocol (scoring against raw gold chains, so DKG agents gain no vocabulary advantage), the DKG yields a +32pp action-accuracy lift over an unstructured baseline (73% vs. 41%). A five-condition ablation reveals that the cause menu and frequency priors alone&amp;#8212;without SQL&amp;#8212;account for the dominant share of this gain (78&amp;#8211;86% of the total lift), establishing that the evaluation infrastructure itself is the primary source of agent improvement. Pre-validated SQL adds a directionally positive but non-significant benefit at n=100, bounded by data sparsity rather than a method ceiling.</description>
      <pubDate>Thu, 16 Jul 2026 15:20:14 GMT</pubDate>
      <guid>https://www.amazon.science/publications/diagnostic-knowledge-graphs-automated-benchmark-construction-and-deterministic-evaluation-for-multi-step-reasoning-agents</guid>
    </item>
    <item>
      <title>Bridging language models and knowledge graphs with controlled natural languages</title>
      <link>https://www.amazon.science/publications/bridging-language-models-and-knowledge-graphs-with-controlled-natural-languages</link>
      <description>Knowledge graphs provide a source of up-to-date structured knowledge, which makes them an ideal counterpart to LLMs. LLMs, by themselves, are not trained to run structured queries internally and can become stale without a source of up-to-date information. We hypothesize that knowledge graphs can be effectively connected to large language models via controlled natural languages. Unlike standard formal query languages, controlled natural languages (CNLs) offer a syntax close to human language. Yet, can be unambiguously converted into formal languages such as SPARQL. In this article, we explore the premise that the extensive pre-training of LLMs on diverse textual data enables them to perform semantic parsing into controlled natural languages more accurately than parsing directly into formal query languages. To evaluate our hypothesis, we constructed a dataset facilitating the comparison between a standard formal language and two controlled natural languages. Our findings show a significant accuracy improvement when using the same amount of controlled natural language training samples. Additionally, fewer samples are required to achieve a desired performance when using CNLs compared to standard query languages. The higher data efficiency of CNLs is particularly important to reduce the complexity and cost of the collection and curation. This enables a more efficient way for LLMs to query KGs.</description>
      <pubDate>Fri, 10 Jul 2026 15:01:49 GMT</pubDate>
      <guid>https://www.amazon.science/publications/bridging-language-models-and-knowledge-graphs-with-controlled-natural-languages</guid>
    </item>
    <item>
      <title>RECoRD: A multi-agent LLM framework for reverse engineering codebase to causal relational diagram</title>
      <link>https://www.amazon.science/publications/record-a-multi-agent-llm-framework-for-reverse-engineering-codebase-to-causal-relational-diagram</link>
      <description>Understanding the behavior and logical structure of complex algorithms is a fundamental challenge in industrial systems. Recent advancements in large language models (LLMs) have demonstrated remarkable code understanding capabilities. However, their potential for reverse engineering algorithms into interpretable causal structures remains unexplored. In this work, we develop a multi-agent framework, RECoRD, that leverages LLMs to Reverse Engineering Codebase to Causal Relational Diagram. RECoRD uses reinforcement fine-tuning (RFT) to enhance the reasoning accuracy of the relation extraction agent. Fine-tuning on expert-curated causal graphs allows smaller specialized models to outperform larger foundation models on domain-specific tasks. Experiments on three real-world use cases - News Vendor, MiniSCOT, and Black-Scholes - demonstrate the effectiveness of our approach. The RFT-trained models significantly outperformed their foundation counterparts, improving F1 score from 0.69 to 0.97 on MiniSCOT. RECoRD also exhibited strong generalization, with models fine-tuned on one use case improving performance on others. We further show how the extracted causal graphs can be leveraged to build a deep-dive assistant that reasons like domain experts, enabling rapid root cause analysis in complex software systems. By automating the construction of interpretable causal models from code, RECoRD has wide-ranging applications in areas such as software debugging, operational optimization, and risk management.</description>
      <pubDate>Tue, 16 Jun 2026 15:54:44 GMT</pubDate>
      <guid>https://www.amazon.science/publications/record-a-multi-agent-llm-framework-for-reverse-engineering-codebase-to-causal-relational-diagram</guid>
    </item>
    <item>
      <title>AutoClimDS: Climate data science agentic AI &amp;#8212; A knowledge graph is all you need</title>
      <link>https://www.amazon.science/publications/autoclimds-climate-data-science-agentic-ai-a-knowledge-graph-is-all-you-need</link>
      <description>Climate data science faces persistent barriers stemming from the fragmented nature of data sources, heterogeneous formats, and the steep technical expertise required to identify, acquire, and process datasets. These challenges limit participation, slow discovery, and reduce the reproducibility of scientific workflows. In this paper, we present a proof of concept for addressing these barriers through the integration of a curated knowledge graph (KG) with AI agents designed for cloud-native scientific workflows. The KG provides a unifying layer that organizes datasets, tools, and workflows, while AI agents&amp;#8212;powered by generative AI services&amp;#8212;enable natural language interaction, automated data access, and streamlined analysis. Together, these components drastically lower the technical threshold for engaging in climate data science, enabling non-specialist users to identify and analyze relevant datasets. By leveraging existing cloud-ready API data portals, we demonstrate that &amp;apos;a knowledge graph is all you need&amp;apos; to unlock scalable and agentic workflows for scientific inquiry. The open-source design of our system further supports community contributions, ensuring that the KG and associated tools can evolve as a shared commons. Our results illustrate a pathway toward democratizing access to climate data and establishing a reproducible, extensible framework for human&amp;#8211;AI collaboration in scientific research.</description>
      <pubDate>Fri, 12 Jun 2026 12:40:31 GMT</pubDate>
      <guid>https://www.amazon.science/publications/autoclimds-climate-data-science-agentic-ai-a-knowledge-graph-is-all-you-need</guid>
    </item>
    <item>
      <title>Linking knowledge to care: Knowledge graph-augmented medical follow-up question generation</title>
      <link>https://www.amazon.science/publications/linking-knowledge-to-care-knowledge-graph-augmented-medical-follow-up-question-generation</link>
      <description>Clinical diagnosis is time-consuming, requiring intensive interactions between patients and medical professionals. While large language models (LLMs) could ease the pre-diagnostic workload, their limited domain knowledge hinders effective medical question generation. We introduce a Knowledge Graph-augmented LLM with active in-context learning to generate relevant and important follow-up questions, KG-Followup, serving as a critical module for the pre-diagnostic assessment. The structured medical domain knowledge graph serves as a seamless patch-up to provide professional domain expertise upon which the LLM can reason. Experiments demonstrate that KG-Followup outperforms state-of-the-art methods by 5% - 8% on relevant benchmarks in recall.</description>
      <pubDate>Thu, 11 Jun 2026 18:09:51 GMT</pubDate>
      <guid>https://www.amazon.science/publications/linking-knowledge-to-care-knowledge-graph-augmented-medical-follow-up-question-generation</guid>
    </item>
    <item>
      <title>From unstructured to structured: LLM-guided attribute graphs for entity search and ranking</title>
      <link>https://www.amazon.science/publications/from-unstructured-to-structured-llm-guided-attribute-graphs-for-entity-search-and-ranking</link>
      <description>Entity search, i.e., finding the most similar entities to a query entity, faces unique challenges in e-commerce, where product similarity varies across categories and contexts. Traditional embedding-based approaches often struggle to capture nuanced context-specific attribute relevance. In this paper, we present a two-stage approach combining Large Language Model (LLM)-driven attribute graph construction with graph-aware LLM ranking. In the offline stage, we extract structured product attributes from unstructured text, and construct a reusable attribute graph with category-aware schemas. In the online stage, we rank retrieved candidates by reasoning over this structured representation rather than raw text, reducing per-product token usage by 57% while improving ranking precision. Experiments show that our approach outperforms multiple baselines under zero-shot scenarios, achieving a over 5% improvement in average precision without requiring training data, generalizes robustly across diverse product categories, and shows immense potential for real-world deployment.</description>
      <pubDate>Wed, 10 Jun 2026 14:24:49 GMT</pubDate>
      <guid>https://www.amazon.science/publications/from-unstructured-to-structured-llm-guided-attribute-graphs-for-entity-search-and-ranking</guid>
    </item>
    <item>
      <title>APEX-MEM: Agentic semi-structured memory with temporal reasoning for long-term conversational AI</title>
      <link>https://www.amazon.science/publications/apex-mem-agentic-semi-structured-memory-with-temporal-reasoning-for-long-term-conversational-ai</link>
      <description>Large language models still struggle with reliable long-term conversational memory: simply enlarging context windows or applying na&amp;#239;ve retrieval often introduces noise and destabilizes responses. We present APEX-MEM, a conversational memory system that combines three key innovations: (1) a property graph which uses domain-agnostic ontology to structure conversations as temporally grounded events in an entity-centric framework, (2) append-only storage that preserves the full temporal evolution of information, and (3) a multi-tool retrieval agent that understands and resolves conflicting or evolving information at query time, producing a compact and contextually relevant memory summary. This retrieval-time resolution preserves the full interaction history while suppressing irrelevant details. APEX-MEM achieves 88.88% accuracy on LOCOMO&amp;apos;s Question Answering task and 86.2% on LongMemEval, outperforming state-of-the-art session-aware approaches and demonstrating that structured property graphs enable more temporally coherent long-term conversational reasoning.</description>
      <pubDate>Mon, 08 Jun 2026 14:56:11 GMT</pubDate>
      <guid>https://www.amazon.science/publications/apex-mem-agentic-semi-structured-memory-with-temporal-reasoning-for-long-term-conversational-ai</guid>
    </item>
  </channel>
</rss>
