← ATH

TOOL · GLOBAL

ESQ-Bench: enterprise SQL benchmark exposes NL2SQL fragility on real schemas

New arXiv benchmark (ESQ-Bench) shows state-of-the-art NL2SQL models (89% on academic tests) fail on actual enterprise SQL dialects and complex schemas. Reveals silent semantic divergence—queries appear correct but return wrong data.

WHY IT MATTERS

Banks deploying LLM-driven SQL agents for data warehouse access face execution risk: model may silently generate syntactically correct but semantically wrong queries, breaking compliance reports and risk calcs without obvious error signals.

Source: arXiv · 2026-08-27

← BACK TO TODAY'S DECK

ESQ-Bench: enterprise SQL benchmark exposes NL2SQL fragility on real schemas — ath — AITechHive