Building Reliable Data Analytics Agents: Lessons from the KDD Cup

NVIDIA's KGMON team took second place in the KDD Cup 2026 Data Agents competition, according to a post on the NVIDIA Developer Blog. The competition required agents to answer natural-language questions across heterogeneous sources including databases, CSV and JSON files, prose documents, PDFs, and briefing videos. Teams had to work with a small, fixed LLM, which made the surrounding harness rather than the model itself the main optimization surface.
The team's core design decision was to collapse structured data access into a single surface. CSV and JSON files were converted into tables inside the existing SQLite database, and a persistent Python environment exposed just two built-in functions for structured work: schema() for inspecting tables and columns, and sql(query) for querying. A preflight schema-scouting step then briefed the agent on tables, likely join keys, duplicate names, look-alike fields, units, null patterns, and row-grain issues before the main reasoning loop started.
The toolset was kept small and opinionated: schema(), sql(query), write_answer(df), and prose_helper(). The write_answer helper was the single approved path for producing final answers, and middleware repaired malformed tool calls so one bad call did not end an attempt. Persistent state retained variables between calls, letting the agent reuse intermediate results.
Documents were treated as a first-class input but kept separate from structured analysis. Direct reads of whole files through open() or .read() were blocked; instead the agent used limited previews and regex searches, then called prose_helper to pass chunks to a separate zero-temperature LLM call with reasoning disabled. That helper operated in an answer mode for rules and thresholds and a table mode for extracting repeating records into a SQL table. Video was preprocessed before the agent loop via keyframe extraction, audio transcription, and transcript-frame alignment; at most one video appeared per task. Traces logged prompts, tool calls, SQL queries, intermediate results, errors, repairs, and final answers, and a specialized inspector agent categorized failures.
Why it matters: The write-up frames harness design, not model capability, as the main lever for reliability, and says the techniques are especially useful when building systems around smaller open models. Several caveats are stated: table extraction from documents suited competition tasks, repeated attempts increase token usage, latency, and compute cost, and autonomous improvement loops can overfit to a benchmark.
Based on reporting from the original publisher. Visit the source for full context and later updates.
Publisher excerpt
The NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a simple idea of making an agent's harness smaller,...