[{"content":"The earlier posts in this series circled the optimizer\u0026rsquo;s execution axis: query rewriting (just asking an LLM to rewrite SQL does almost nothing, the blind spot of rule-based rewriting on DSB), plan tuning (treating the LLM as a plan-tuner rather than an optimizer), and physical design (when the index tuner\u0026rsquo;s cost model lies). This one drills down to the lowest input the optimizer has — cardinality estimation (CardEst).\nThe source is an empirical study from Peking University and ByteDance, released in late March 2026 (arXiv:2603.28080), with a title posed as an honest question: Can Large Language Models be a Cardinality Estimator? An Empirical Study. It doesn\u0026rsquo;t presume the answer; instead it drags Llama-3 8B into PostgreSQL\u0026rsquo;s optimizer and carefully measures both the accuracy and the end-to-end bill.\nThe conclusion comes in two halves, and it\u0026rsquo;s the two halves together that make this paper worth reading: on accuracy the LLM wins cleanly, end-to-end it loses badly — at first.\nWhy cardinality estimation is the optimizer\u0026rsquo;s vital organ The optimizer picks an execution plan via the cost model; whether the cost model is right rests almost entirely on cardinality estimation — how many rows each sub-query (especially after multi-table joins) will actually produce. Over-estimate and the optimizer thinks it should build a hash join on a huge table; under-estimate and it thinks a few rows can be nested-loop\u0026rsquo;d freely. A cardinality off by an order of magnitude can knock the downstream join order, join algorithm, and parallelism all askew, and the plan degrades outright.\nTraditional estimators rely either on histograms plus independence assumptions (PostgreSQL\u0026rsquo;s default) or on learned models (MSCN, DeepDB, NeuroCard, and the newer PRICE). Their shared headache is correlation under multi-table joins: the more joins, the more distorted the distribution, the wilder the estimate. That is exactly the terrain this paper wants to see whether an LLM can handle.\nEvidence one: on accuracy, the LLM genuinely wins On method, the team did not cram the schema into some MLP encoder — they tried, and it overfit badly. The final recipe is a triple: a carefully designed prompt (fed coarse-grained statistics, and even the outputs of other estimators like PRICE stuffed into the context), LoRA fine-tuning (the key being to generate the number digit by digit as tokens rather than regress a scalar), and a self-correction loop at inference time. The default model is Llama-3 8B, the underlying engine PostgreSQL, with PilotScope injecting the estimated cardinalities back into the optimizer.\nAccuracy is measured by Q-error (the ratio of estimate to truth, taken in the larger direction; 1 is perfect). One setting captures the point best: train only on queries with fewer than 3 joins, test entirely on queries with more than 3 joins — forcing the model to extrapolate to harder correlations. The 99th-percentile Q-error column:\nEstimator 99th-pct Q-error PostgreSQL 52594 MSCN 17239 DeepDB 20236 PRICE (best baseline) 10672 Llama-3 + fine-tuning 2581.73 The gap isn\u0026rsquo;t a few percent, it\u0026rsquo;s an order-of-magnitude tightening of the tail. On this IMDB table, the 99th percentile tightens from PRICE\u0026rsquo;s 10672.53 to 2581.73 — a 75.81% reduction. On the standard STATS setting, the 99th-pct Q-error drops by up to 74.1% relative to the best baseline. Across IMDB, STATS, ErgastF1, and Genome, the LLM beats the strongest baseline PRICE on nearly every setting. A tight tail means the very queries most likely to drive the optimizer off a cliff get rescued.\nIf the paper stopped here, it\u0026rsquo;d be another \u0026ldquo;the LLM won again\u0026rdquo; piece. But the authors went on to measure end-to-end.\nEvidence two: end-to-end, it backfires — at first Feed more accurate cardinalities back into the optimizer, the plan should be better, the query faster — that\u0026rsquo;s the intuition. The measurements flipped on several workloads.\nOn JOB-light and ErgastF1, plain Llama+PFT\u0026rsquo;s end-to-end total time actually exceeds the strongest baseline PRICE: it estimates more accurately, and on query-execution time alone it beats PG, PRICE, and ASM by a wide margin — but every estimate requires one inference pass of an 8B model, and that latency, accumulated into the total, eats up all the plan gain the accuracy bought on these two workloads. \u0026ldquo;More accurate\u0026rdquo; has, for the first time here, decoupled from \u0026ldquo;faster.\u0026rdquo;\nThis is the crux of the whole paper, and the one sentence the perf lens should keep: accuracy is the benefit, the cost of one estimate is the price, and end-to-end is the two subtracted. On the accuracy board alone, the LLM sweeps; count the wall-clock of every inference, and the bill changes instantly.\nIt doesn\u0026rsquo;t lose everywhere, of course. On more complex, larger-plan-space workloads like STATS and Genome, the plan improvement from the accuracy advantage far outweighs that inference overhead, and end-to-end still nets out ahead. The question shifts from \u0026ldquo;can the LLM do it\u0026rdquo; to \u0026ldquo;when is calling the LLM worth it.\u0026rdquo;\nEvidence three: cost-aware gating, a very old wisdom The authors\u0026rsquo; fix isn\u0026rsquo;t in the model, it\u0026rsquo;s in the scheduling — and it\u0026rsquo;s an old trick anyone from query engines recognizes on sight: use the cost model as a bouncer.\nConcretely (the paper calls it LLMs+PFT+cost): for each sub-query, first use PostgreSQL\u0026rsquo;s own cost model to estimate its execution cost; only the high-cost sub-queries above a threshold get to spend on an LLM call to refine the cardinality; very low-cost sub-queries just use PostgreSQL\u0026rsquo;s native estimate. The threshold is calibrated from the empirical relationship between \u0026ldquo;estimated cost vs actual runtime.\u0026rdquo;\nThe logic is clean: high-cost sub-queries tend to have many joins and strong correlation — precisely where the baselines degrade worst and the LLM\u0026rsquo;s edge is largest, so spending one inference here to buy back a plan that doesn\u0026rsquo;t degrade several-fold is worth it; low-cost sub-queries are already estimated accurately enough by the baseline, and calling the LLM is pure wasted latency. The gate spends the expensive compute exactly on the sub-queries where it can change the outcome.\nThis is the same sentence this series keeps returning to: the LLM doesn\u0026rsquo;t replace the optimizer, it\u0026rsquo;s used under the optimizer\u0026rsquo;s discipline. Last post, the cost model was the thing the LLM patched the blind spots of; this post, it flips into the gate that constrains the LLM — the cost model is on both ends.\nEvidence four: laying the cost bill out flat Worth being clear about the cost side too, so accuracy isn\u0026rsquo;t all that\u0026rsquo;s remembered.\nPre-fine-tuning (PFT): one-time, offline, about 12 hours, but the product is reusable across any database instance, so amortized it\u0026rsquo;s cheap. Target-DB fine-tuning: LoRA, about 50–60 minutes, on par with most baselines, and even shorter than DeepDB, FactorJoin, and ALECE. Single-inference latency: the structural weak point, and the root of the end-to-end backfire in evidence two; the gate exists for it. Self-correction: iterative correction reduces hallucination, but the authors measured that on IMDB 85.03% of queries are accurate enough with zero iterations, and only 9.24% need up to 5 rounds — the bulk doesn\u0026rsquo;t actually iterate, and the overhead is less dramatic than feared. The hardware is 8× A-800 (80GB), 512GB RAM, trained with LLaMA-Factory. The code is open-source (PKU-SDS-lab/EA-LLM-IN-DB). The scale isn\u0026rsquo;t small, but the training-side bill amortizes; what truly bottlenecks is always that one inference pass.\nClosing: the cost model as bouncer String the four pieces of evidence together and the paper answers its title\u0026rsquo;s question, with a caveat: a large model can be a cardinality estimator, provided you don\u0026rsquo;t call it for every sub-query.\nOn accuracy it genuinely won, and won on the hardest multi-join tail; but amortize inference latency into end-to-end and calling it blindly backfires. The real engineering contribution isn\u0026rsquo;t the fine-tuned Llama, it\u0026rsquo;s that plain gate — letting the optimizer\u0026rsquo;s own cost model decide which sub-queries deserve an LLM inference. Spending expensive compute on the cutting edge isn\u0026rsquo;t unique to CardEst; it\u0026rsquo;s most likely the general pattern for LLMs entering every stage of the query engine: the model is responsible for being more accurate where it\u0026rsquo;s hard, the cost model for constraining it to be called where it pays.\nThe next thing that should be examined this way is which stage left in the optimizer can both withstand the LLM\u0026rsquo;s accuracy dividend and bear the latency bill of every call.\nSources arXiv:2603.28080 — Can Large Language Models be a Cardinality Estimator? An Empirical Study, Peking University + ByteDance + UPenn, 2026-03-31. Code: https://github.com/PKU-SDS-lab/EA-LLM-IN-DB Accessed: 2026-05-31. ","permalink":"https://yaooqinn.github.io/posts/query-engines/llm-cardinality-estimator/","summary":"Cardinality estimation is the heart of the optimizer. A team from Peking University and ByteDance fine-tunes Llama-3 8B to do CardEst, and on workloads like IMDB and STATS the 99th-percentile Q-error drops by up to 74.1% versus the strongest baseline (PRICE) — the accuracy win is real. But end-to-end, it backfires: on JOB-light and ErgastF1 the LLM\u0026rsquo;s more accurate plans are dragged down by its own inference latency, with total time exceeding even the strongest baseline PRICE. The real engineering contribution isn\u0026rsquo;t the model — it\u0026rsquo;s the gate that uses the optimizer\u0026rsquo;s own cost model as a bouncer: call the LLM only for high-cost sub-queries, leave the rest to the old methods.","title":"LLMs as Cardinality Estimators: Accurate, But Only If You Don't Call Them Every Time"},{"content":"The earlier posts in this series were all about query rewriting: just asking an LLM to rewrite SQL does almost nothing, the blind spot of rule-based rewriting on DSB, and treating the LLM as a plan-tuner rather than an optimizer. This one switches axes — to physical design, specifically index tuning.\nThe source is an evaluation paper from Microsoft\u0026rsquo;s own team (Xiaoying Wang, Wentao Wu, Vivek Narasayya, Surajit Chaudhuri) released in March 2026 (arXiv:2603.09181), with a deliberately restrained title: Evaluating the Practical Effectiveness of LLM-Driven Index Tuning with Microsoft Database Tuning Advisor. Narasayya and Chaudhuri have been the foundational figures behind SQL Server index tuning (AutoAdmin / DTA) for over two decades — them stepping in to evaluate LLMs is itself worth reading.\nAnd this paper has a particularly honest posture: it\u0026rsquo;s not here to announce that LLMs won. It\u0026rsquo;s here to tell you where the LLM wins, and why it still can\u0026rsquo;t go to production.\nOne query, two opposite endings Start with the single most striking number in the paper.\nOn a real enterprise customer workload called Real-R, there\u0026rsquo;s a query 22. Microsoft\u0026rsquo;s DTA — the industry\u0026rsquo;s SOTA commercial index tuner — recommends a set of indexes for it, and those indexes regress the query\u0026rsquo;s execution time by nearly 10x (a severe query performance regression, or QPR).\nSame query, hand it to GPT-5: execution time drops from 10 seconds to 4 seconds, roughly a 60% improvement.\nSame query optimizer, same data, same SQL. One pushed it off a cliff; the other walked around the edge. Why?\nWhere the blind spot comes from: the what-if cost model To understand this, you first need to know how DTA works.\nDTA is a textbook cost-based architecture with three stages: (1) workload analysis, identifying indexable columns from the queries (columns appearing in filters and join conditions); (2) candidate index generation, deciding the key columns, included columns, and column ordering of each potentially useful index; and (3) configuration enumeration, selecting a subset of candidates that minimizes the estimated execution cost of the whole workload while respecting constraints like the maximum number of indexes or storage budget.\nThe operative word is estimated. DTA estimates each configuration\u0026rsquo;s cost using the optimizer\u0026rsquo;s what-if API — which can answer \u0026ldquo;if this index existed, roughly how much would this query cost?\u0026rdquo; without actually materializing the index. It\u0026rsquo;s a clever design that avoids the expensive cost of building indexes.\nBut it inherits the original sin of the cost model: when the estimate is wrong, the recommendation is suboptimal, and in extreme cases it causes a production incident. This isn\u0026rsquo;t a bug in DTA\u0026rsquo;s implementation; it\u0026rsquo;s the twenty-year-old problem of cost-based architectures — cardinality estimates propagate errors, join result sizes get mis-estimated, exactly as Leis et al.\u0026rsquo;s well-known How Good Are Query Optimizers, Really? demonstrated repeatedly. The 10x regression on query 22 is what-if estimating a genuinely bad set of indexes as a bargain.\nThe LLM route bypasses this whole explicit cost model. It doesn\u0026rsquo;t compute cost; it draws directly on the web-scale-trained intuition for \u0026ldquo;what kind of query wants what kind of index\u0026rdquo; and hands you a configuration. Where DTA was misled by its cost estimate, that intuition happened not to fall into the same hole.\nWhy this evaluation is credible There are plenty of papers on LLMs for database tuning, and most share a flaw: they benchmark on public datasets that have almost certainly leaked into the LLM\u0026rsquo;s training set. This paper deliberately avoids that:\nReal execution time, not estimated cost. Each query is actually run 5 times and the median is taken, with a 300-second timeout. It measures wall-clock, not the cost the optimizer reports about itself — the same methodological stance as #1 in this series. Four of the five workloads are real enterprise customer workloads (Real-D / Real-M / Real-R / Real-S), with CTEs, views, and many manually created indexes — the \u0026ldquo;dirt\u0026rdquo; that public benchmarks (here only TPC-H SF10 was used) lack. The baseline is the real commercial DTA, not the simplified index recommender common in papers. The primary LLM is GPT-5 (they also tried DeepSeek-R1, Qwen3, and GPT-4o; GPT-5 was best, and all reported results use it). This setup makes the conclusions far more solid than \u0026ldquo;yet another leaderboard paper.\u0026rdquo;\nComplementarity: where the LLM wins In the single-query setting, within 5 invocations, GPT-5 matches or beats DTA in most cases. And there\u0026rsquo;s a counterintuitive detail: it often achieves the same effect with fewer indexes. On query 4 of Real-D, DTA recommends 17 indexes; GPT-5 recommends only 7, and most of those 7 are actually used by the execution plan.\nMore important is where the LLM wins. The places where it substantially beats DTA are exactly where DTA was led astray by inaccurate cost estimates — on query 4 of Real-D and query 27 of Real-M, DTA\u0026rsquo;s recommendations both caused QPRs, and the LLM avoided them. The paper puts it bluntly: those most useful LLM-recommended indexes were never even candidates in DTA\u0026rsquo;s candidate-generation stage — because what-if judged them to have higher estimated cost and culled them early.\nThe paper also does something interesting: it distills the rules of thumb from GPT-5\u0026rsquo;s reasoning — prioritize indexes that cut expensive scans, order key columns by cues from the plan\u0026rsquo;s filters/joins/sorts, build covering indexes where possible (to enable index-only scans), and ignore small-table scans. Then it writes a purely rule-based tuner that calls no LLM from those rules. In other words, part of the LLM\u0026rsquo;s value here can be distilled into deterministic rules.\nThree walls: why it can\u0026rsquo;t go to production yet If the paper stopped here, it would be an LLM puff piece. But the Microsoft team honestly raises three walls.\nWall one: high variance. Invoke GPT-5 five times on the same query and the gap between best and worst is enormous. On TPC-H\u0026rsquo;s query 20, the worst case degrades into an outright QPR. If you take the worst-case result for each query, the LLM no longer wins a single query, and in most cases performs substantially worse than DTA. That beautiful 10s→4s ending could be a disaster on the next invocation. The intuition is unstable.\nWall two: direct integration degrades. A natural idea: since the LLM can come up with good indexes that aren\u0026rsquo;t in DTA\u0026rsquo;s candidate pool, why not inject the LLM\u0026rsquo;s recommendations into DTA\u0026rsquo;s candidate pool and let DTA pick? Expanding the search space should only help. The opposite happened — after merging the single-query LLM recommendations into the candidate pool, the final configuration got significantly slower. The cause, again, is what-if: the pool is larger, but the selection criterion is still that inaccurate cost estimate, so it picks a worse configuration from the bigger pool. (The one exception: workload-level multi-query recommendations can improve DTA by more than 2x — but that, conversely, shows the problem is in the cost model, not the candidate indexes.)\nWall three: validation is prohibitively expensive. So step back: since the estimate can\u0026rsquo;t be trusted, why not just materialize the candidate indexes, run the workload, and pick the best by real execution time? The paper does an end-to-end time breakdown, and the answer is: validation costs far more than tuning itself, and the bottleneck is index creation. Actually materializing the candidate indexes and re-running the workload — index creation alone accounts for a disproportionately large share of the total cost. And the diversity of LLM recommendations amplifies this — different invocations give different configurations, and the set of indexes to validate keeps piling up. In production, this road is essentially a dead end.\nWhat it means for engine builders Put the three walls together and the conclusion is clear — and it dovetails exactly with the through-line of #1 in this series: bypassing the cost model\u0026rsquo;s approximations is not free; the price is reliability.\nIn #1, the LLM rewriting SQL bypassed the cost model using real execute() wall-clock, at the cost of needing training and feedback. In this paper, the LLM recommending indexes bypasses what-if, at the cost of high variance, non-integrability, and un-cheap validation. Same exchange rate, different underlying.\nSo the correct positioning of the LLM in index tuning today is not as a replacement for DTA, but as a source of the candidate indexes the cost model can\u0026rsquo;t see. The real open problem is not \u0026ldquo;can the LLM come up with good indexes\u0026rdquo; — it can; query 22 is the proof — but \u0026ldquo;how do you cheaply trust the candidates it offers, without rebuilding indexes across the entire warehouse to find out?\u0026rdquo;\nFor engine builders, that\u0026rsquo;s a very concrete architectural signal. Physical-design automation — index recommendation and materialized-view recommendation in systems like Spark and Kyuubi — should next be thinking not about \u0026ldquo;should we switch to an LLM,\u0026rdquo; but about how to decouple \u0026ldquo;generating candidates\u0026rdquo; from \u0026ldquo;trusting candidates\u0026rdquo;: let the LLM (or any cost-model-free source) expand the candidate space, and treat \u0026ldquo;how to build trust cheaply in a world where what-if is inaccurate\u0026rdquo; as a separate sub-problem worth investing in on its own. This paper doesn\u0026rsquo;t give the answer, but it asks the question correctly.\n","permalink":"https://yaooqinn.github.io/posts/query-engines/llm-index-tuning-vs-dta/","summary":"A Microsoft team evaluates LLM-driven index tuning on real enterprise customer workloads. On query 22 of Real-R, the SOTA commercial tuner DTA recommends indexes that cause a near-10x regression; on the same query, GPT-5 cuts execution time from 10 seconds to 4. The LLM wins precisely where the what-if cost model is wrong. But that intuition is high-variance, can\u0026rsquo;t be bolted into the existing architecture, and can\u0026rsquo;t be validated cheaply — it\u0026rsquo;s not a replacement for DTA today, it\u0026rsquo;s a source of the candidate indexes DTA can\u0026rsquo;t see.","title":"When the Index Tuner's Cost Model Lies: Where LLMs See What DTA Can't"},{"content":"My previous post argued that just asking an LLM to rewrite SQL does almost nothing — without plan signals or feedback, prompting a commercial LLM on TPC-H 10GB takes 78.81s down to 74.92s, statistically indistinguishable from doing nothing.\nThat post left an obvious follow-up. What about the traditional path — the rule-based and learned rewriters that the database research community has been refining for decades? Surely those are stable?\nAfter reading the QUITE paper\u0026rsquo;s Table 3 (arXiv:2506.07675, June 2025), my answer to that question changed: they\u0026rsquo;re not stable. They only look stable on the benchmarks they were trained on.\nOne number that collapses when you change the schema QUITE compares rewriters across three benchmarks. Pull out the row for LearnedRewrite (SIGMOD'22, \u0026ldquo;LR\u0026rdquo; hereafter):\nBenchmark Original mean latency LR rewritten Improvement TPC-H (SF=10) 69.84s 37.57s −46.2% DSB (SF=10) 32.62s 31.93s −2.1% Calcite 23.88s 22.88s −4.2% (Source: QUITE paper Table 3. All benchmarks at SF=10, PostgreSQL, three-run average, 300s timeout.)\nLR\u0026rsquo;s TPC-H report card is respectable — nearly half the execution time gone. If you only looked at TPC-H, you\u0026rsquo;d conclude this is a real rewriter.\nBut the same LR, same code, on DSB takes 32.62 to 31.93 — a 2.1% gain. Inside benchmark noise. Effectively no rewrite happened.\nDSB isn\u0026rsquo;t an obscure dataset — it\u0026rsquo;s Microsoft\u0026rsquo;s 2021 decision-support benchmark, with deeper nested subqueries, more CTEs, and more realistic filter patterns than TPC-H. The rewrite headroom is objectively there: in the same paper QUITE takes DSB from 31.93 down to 5.85. The equivalent-rewrite space exists; LR just can\u0026rsquo;t see it.\nThe paper says so itself The most direct statement of the problem is in QUITE\u0026rsquo;s Section 7.2:\n\u0026ldquo;The LR we use is trained on the TPC-H, LR can proficiently identify effective rewrite rules and navigate their constrained search spaces.\u0026rdquo;\nLR works on TPC-H because it was trained on TPC-H.\nThe phrasing is mild, but it surfaces a structural issue with three decades of rule-rewriter research. Whether \u0026ldquo;training\u0026rdquo; is meant in the learning sense or not, rule libraries themselves are shaped by benchmark distributions:\nWhich rules get into Calcite/Orca\u0026rsquo;s rule library? Usually because a paper demonstrated their value on a benchmark. Which rules get heavily refined? Usually because they show measurable gains on TPC-H / TPC-DS / JOB. Which rules get pruned? Ones that contribute nothing — or regress — on the standard benchmarks. The rule set evolved against the academic benchmark distribution, not the real-world query distribution. Rule-based rewriting is pattern matching at its core; if a pattern isn\u0026rsquo;t in the training distribution, the system can\u0026rsquo;t see it. Not my words — QUITE quotes Leis et al. (VLDB'15, the canonical join-order benchmark paper):\n\u0026ldquo;fixed rewrite rules rely on pattern matching and therefore are fundamentally unable to optimize unseen or complex query patterns\u0026rdquo;\nWhat \u0026ldquo;rule-based systems are reliable\u0026rdquo; actually means Back to those two numbers, −46% and −2.1%.\nIf your impression of rule-based rewriters is \u0026ldquo;mature, proven, stable,\u0026rdquo; that\u0026rsquo;s because almost every published evaluation you\u0026rsquo;ve seen runs on TPC-H, TPC-DS, or JOB. These benchmarks are the training set for rule-based systems. Looking good on the training set isn\u0026rsquo;t \u0026ldquo;stable and reliable.\u0026rdquo;\nDSB isn\u0026rsquo;t part of that training set — it came out after LR was published, and the rule library hasn\u0026rsquo;t had time to grow the patterns it needs. Result: −2.1%.\nThe general form of this problem: there is virtually no public data on how rule-based rewriters perform on real production workloads. What does production SQL actually look like at a real company? Nesting depth, predicate shapes, join patterns, UDF usage — almost none of it resembles TPC-H. When you deploy a rewriter that won \u0026ldquo;−46% on TPC-H\u0026rdquo; against a real workload, you might get −5%, or 0, or some queries getting slower (rule fires, payoff is negative).\nThis is why the LLM path is still interesting for rewriting, despite the previous post\u0026rsquo;s bleak baseline. Rule-system search space is closed, hand-curated, and lags behind real workload evolution. LLM pattern space is open and can extrapolate to unseen schema shapes. The former has a hard ceiling inside its training distribution; the latter has an unclear ceiling but at least won\u0026rsquo;t freeze when it encounters DSB for the first time.\nQUITE\u0026rsquo;s own numbers on the three benchmarks — bringing LR\u0026rsquo;s −46% / −2% / −4% up to roughly −63% / −82% / −58% — make the shape of the blind spot visible: small relative gains on TPC-H, huge relative gains on DSB. That\u0026rsquo;s exactly what a benchmark-overfit baseline looks like: little headroom left where it was trained, lots of headroom left everywhere else.\nThree questions still open The thesis is clear: rule-based rewriters aren\u0026rsquo;t broken — they\u0026rsquo;re just untested outside their benchmark. But QUITE\u0026rsquo;s recipe still leaves several open problems:\n1. How do you measure a rewriter\u0026rsquo;s coverage generalization?\nAll current benchmarks are single points — one number on TPC-H, one on DSB, one on Calcite. But generalization isn\u0026rsquo;t about how high any single point is, it\u0026rsquo;s about the variance across points. A system that scores −46% / −2% / −4% and one that scores −20% / −15% / −18% can have the same average but very different stability profiles. The field needs a cross-schema stability metric, not more benchmark data points.\n2. Is there public data on the gap between rewriter training distribution and production distribution?\nVendors with massive production SQL — Databricks, Snowflake, Alibaba, ByteDance — could publish \u0026ldquo;real query shape distribution vs TPC-H shape distribution\u0026rdquo; comparisons: subquery depth histograms, predicate count distributions, join type breakdowns. Right now, the entire industry is operating on intuition about whether rule-based rewriters will land in production.\n3. Where are the cost-quality-coverage boundaries between free LLM generation, FSM-constrained generation, and rule enumeration?\nFree LLM generation has equivalence risk (the previous post noted E³-Rewrite without fine-tuning hit only 84.8% equivalence). Rule enumeration has the coverage blind spots discussed here. QUITE-style \u0026ldquo;FSM frame + LLM generation + Calcite equivalence check\u0026rdquo; sits in the middle. But the actual curves of cost (tokens, latency, ops), quality (equivalence rate, hit rate, average speedup), and coverage (number of schemas where it stays stable) across all three — no public comparison exists. This is the real industrial-readiness gap.\nI don\u0026rsquo;t have an actionable takeaway. What I want to leave is two numbers — −46% and −2.1% — and one caution:\nA rewriter\u0026rsquo;s score on any single benchmark only proves it works on that benchmark. Before debating \u0026ldquo;should we use LLMs to rewrite SQL,\u0026rdquo; it might be worth asking the prior question: our current rule-based rewriter — does it score −46% on our actual workload, or −2%?\nIf it\u0026rsquo;s the latter, then the real question isn\u0026rsquo;t whether LLMs work. It\u0026rsquo;s whether the way we\u0026rsquo;ve been evaluating rewriters this whole time is the problem.\n","permalink":"https://yaooqinn.github.io/posts/query-engines/rule-rewrite-blindspot-dsb/","summary":"On TPC-H 10GB, a state-of-the-art learned rewriter cuts mean execution time from 69.84s to 37.57s — a 46% win. On DSB 10GB, the same rewriter takes 32.62s to 31.93s — a 2.1% non-event. The gap isn\u0026rsquo;t query difficulty; it\u0026rsquo;s whether the benchmark is in the rewriter\u0026rsquo;s training distribution. \u0026ldquo;Rule-based systems are stable and reliable\u0026rdquo; is often a benchmark artifact, not an engineering fact.","title":"−46% or −2%? Rule-Based Rewriters Only Work at Home"},{"content":"\nMy previous two posts (LLM × join order and the blind spot of rule-based rewrites) were both about letting an LLM rewrite SQL text — string in, string out, execution engine unchanged. That route is easy to start with, but it can\u0026rsquo;t touch problems that only show up at the physical-plan level: join order, projection column order, physical operator choice.\nWhat if you let the LLM rewrite the physical plan directly?\nI recently read a February 2026 paper, Making Databases Faster with LLM Evolutionary Sampling (arXiv:2602.10387), with the companion repo BauplanLabs/Making-Databases-Faster-with-LLM-Evolutionary-Sampling. That\u0026rsquo;s exactly what it does. On DataFusion, GPT-5 rewrites physical plans through structural transforms and lands a 4.78× geometric-mean speedup on TPC-H SF10.\nThe number is nice, but the more transferable thing is the prompt. I read sql_optimization_prompts.py line by line. Exactly 120 lines.\nHow those 120 lines are spent Roughly by content:\nLayer What\u0026rsquo;s in it Share L1 · Input schema Defines the three inputs (structure, succinct_table_info, query) and gives a three-node sample JSON ~20% L2 · Methodology Step 1: cardinality estimation. Step 2: apply join-side selection and join reordering on those estimates ~30% L3 · Invariant contract + output schema Three invariants, a projection-index calculation rule with a worked example, output must be a \u0026lt;patch\u0026gt;[...]\u0026lt;/patch\u0026gt; block wrapping an RFC 6902 JSON Patch array ~50% The ratio is what made me stop. Methodology gets 30%. The other 70% is spent on defining what counts as a valid output.\nThat is counter-intuitive. My default assumption was that a prompt for query optimization should mostly teach the LLM how to optimize — what makes a good plan, when to broadcast, when to shuffle. The DBPlanBench authors compress that to 30% and put the remaining 70% into making the LLM\u0026rsquo;s output machine-verifiable.\nLook at the results again — 4.78× geometric mean, and roughly half of LLM outputs pass the validator (paper §4) — and the ratio starts making sense. With a looser prompt, the model would produce plans that look reasonable but don\u0026rsquo;t compile; if only 1 in 5 passes, the evolutionary loop has nothing to search over.\nThree designs worth borrowing ① \u0026ldquo;Be not lazy\u0026rdquo;: explicitly forbid the empty response One block in the prompt reads:\nYou should assume that the current plan can almost always be improved and must not be lazy: actively search for semantics-preserving structural changes instead of defaulting to making no changes. By default, the JSON patch array you output should contain at least one operation that changes the plan structure. Returning an empty array [] ... is acceptable only in critical cases ... That made me smile. When an LLM has to decide whether something can be improved, it tends to take the lazy path and return an empty array. Most prompt authors know this, but few bother to write it down. The authors here flip the default — try is the baseline, empty is the escape hatch.\nThis pattern has nothing to do with query optimization specifically. It applies to any agent-style tool call where you don\u0026rsquo;t want the model to bail out with \u0026ldquo;looks fine to me\u0026rdquo;: PR review, performance triage, error attribution.\n② \u0026ldquo;Use semantics, not defaults\u0026rdquo;: delegate cardinality estimation to world knowledge Step 1 includes:\nCRITICAL: Do not use defaultFilterSelectivity or other default values. Instead, perform semantic analysis of column names, table contexts, and filter predicates to make intelligent cardinality estimates based on real-world knowledge. This is the central bet of the whole approach. Cost-based optimizers fall back to constants when statistics are missing (PostgreSQL\u0026rsquo;s equality default is DEFAULT_EQ_SEL = 0.005, see selfuncs.h). DBPlanBench bets that the LLM can read the column name order_status and the predicate = 'completed' and come up with something like 0.7.\nIt\u0026rsquo;s also the biggest limitation of the method — more on that below.\n③ Projection-index calculation: invariants written as code comments The bit I found most interesting is in L3:\nWhen you swap the left and right inputs of a HashJoin, you MUST update the projection indexes to reflect the new schema order. The projection calculation follows this rule: - If projection references a left field: use the index as-is - If projection references a right field: offset by len(left_schema), i.e., len(left_schema) + right_projection[i] Then four worked examples: A[name, id] × B[dept_name, budget], original projection [0, 3], after swap it should be [2, 1] — spelled out step by step.\nI rarely see this style in prompt engineering: the engine\u0026rsquo;s physical invariants written as code comments to the LLM. The common move is to write \u0026ldquo;preserve semantics after swap\u0026rdquo; and then get back plans that compile but reference columns at the wrong positions. The authors hand the LLM the exact arithmetic. The model doesn\u0026rsquo;t need to understand — it only needs to repeat.\nWhy the 4.78× doesn\u0026rsquo;t transplant cleanly My first reaction was to copy the recipe. My second was to slow down. A real share of that 4.78× is a DataFusion-shaped tailwind, not evidence that LLMs out-think cost models.\nDataFusion\u0026rsquo;s physical optimizer is famously sparse. Statistics propagation only landed in the last few months, histograms are still in RFC stage, and equality-filter selectivity in many paths really does fall back to that 0.005-class default. When the LLM uses semantics to land on 0.7 for order_status = 'completed', the baseline it\u0026rsquo;s beating is 0.005. That ~100× gap is \u0026ldquo;DataFusion lacks statistics\u0026rdquo;, not \u0026ldquo;LLM beats CBO.\u0026rdquo;\nRun the same method on Spark\u0026rsquo;s CBO with histograms, or on Photon or Trino with a real cost model, and the headroom shrinks. Not zero — just much smaller. The 4.78× does not extrapolate.\nThe projection-index rule has the same flavor. DataFusion schemas are flat Vec\u0026lt;Field\u0026gt;, so index arithmetic is clean. Catalyst uses StructType + ExprId + AttributeReference, and \u0026ldquo;how projection moves after a join swap\u0026rdquo; is scattered across transforms like ResolveReferences and BindReferences. There is no single rule to copy-paste into a prompt. The work has to start one step earlier: extract those invariants explicitly — and that itself is a non-trivial engineering project.\nWhat actually transfers Strip the numbers away and three pieces hold up:\nThe three-layer structure (schema / methodology / invariants) and especially the ratio — 70% spent on what counts as a valid output. Not specific to query optimization. It applies to any task that produces a structured artifact the system has to validate. JSON Patch (RFC 6902) as the output format. Industry standard, off-the-shelf appliers, and it matches the diff-shaped reasoning LLMs already do. No need to invent a wire format. The \u0026ldquo;be not lazy\u0026rdquo; instruction. Whenever you don\u0026rsquo;t want the LLM to short-circuit with \u0026ldquo;no change needed\u0026rdquo;, this pattern is worth keeping. The cardinality-via-semantics part — the source of the 4.78× on DataFusion — is more likely a liability on Spark or any engine with real statistics. Invert it: inject the real numbers into the prompt and stop asking the model to guess.\nOne question I\u0026rsquo;m not answering What I haven\u0026rsquo;t decided yet: whether a ~50% validator pass rate is good enough for production.\nDBPlanBench runs offline. Evolutionary sampling can afford to spend a night, sample hundreds of plans, and keep the fastest. 50% pass rate is fine. An online plan tuner is a different game — five rewrites per query before one compiles costs both latency and tokens, and it\u0026rsquo;s not obvious which side runs out of budget first.\nI\u0026rsquo;m not sure how far this route goes yet. Reading the prompt end to end gave me at least a clearer map of where to look next time someone shows up selling \u0026ldquo;let the LLM rewrite plans.\u0026rdquo;\nSources for the numbers in this post\n4.78× geometric mean: arXiv:2602.10387, TPC-H SF10 120-line prompt: wc -l sql_optimization_prompts.py against the 2026-05-26 commit ~50% validator pass rate: paper §4 PostgreSQL DEFAULT_EQ_SEL = 0.005: selfuncs.h JSON Patch: RFC 6902 ","permalink":"https://yaooqinn.github.io/posts/query-engines/prompt-anatomy-for-plan-generation/","summary":"DBPlanBench gets GPT-5 to deliver a 4.78× geometric-mean speedup on DataFusion TPC-H SF10 by letting the model rewrite physical plans directly. I read its sql_optimization_prompts.py end to end — 120 lines, 30 of methodology, 90 of contract. That ratio is the most transferable thing in the paper.","title":"Anatomy of a 120-Line Prompt That Lets an LLM Rewrite Physical Plans"},{"content":"\nMy previous three posts (LLM × join order · the blind spot of rule-based rewrites · anatomy of a 120-line prompt) all asked the same question from different angles: how should an LLM participate in query optimization? Those posts dissected mechanism — rewriting SQL text, rewriting plans, the prompt that drives the rewrite. This one asks a different question: where in the optimizer pipeline does the LLM belong?\nMy answer: behind the optimizer, emitting small patches, not replacing the cost-based optimizer. This is not a claim about technical novelty. It is a claim about engineering governance.\nThree routes, laid out Across the LLM × QO work I\u0026rsquo;ve read this year, the integration paths fall into roughly three buckets:\nRoute Representative work What the LLM emits Output granularity A. Rewrite SQL text (SQL → SQL rewriting, covered in earlier posts) Rewritten SQL string Whole statement B. LLM makes plan decisions directly Databricks 2026-04 blog: LLM agent picking join order1 A specific decision (e.g. join order) Whole-plan dimension C. LLM emits patches against the optimizer\u0026rsquo;s plan DBPlanBench (arXiv:2602.10387)2 RFC 6902–style JSON Patch Local edits A has been discussed. Its core difficulty is that the SQL text layer doesn\u0026rsquo;t carry enough information to fix problems that only show up in the physical plan (join order, projection column order, operator choice). B and C differ in a finer way, which is what this post is about.\nWhy the patch route is worth taking seriously DBPlanBench is the most complete open-source implementation of route C. The workflow: let DataFusion compile SQL to a physical plan using its own optimizer, flatten that plan into JSON, feed it to GPT-5, have the LLM emit RFC 69023 JSON patches (e.g. swap a hash join\u0026rsquo;s build/probe sides), and apply the patches back to DataFusion for real execution.\nOn TPC-H and TPC-DS (SF=3/10), the paper reports up to 4.78× speedup on individual queries2. That number is not the point of this post — I want to surface its caveats first:\nThe baseline is DataFusion\u0026rsquo;s own optimizer; there is no head-to-head against learned optimizers (Bao/Balsa) or against Photon / Spark CBO. The experiments stop at SF=10 (\u0026lt;10 GB), not TB. \u0026ldquo;Transfer to larger scale\u0026rdquo; is demonstrated only between SF=3 and SF=10, via a one-shot deterministic rewrite script — not a true TB-scale validation. So I\u0026rsquo;d read 4.78× as \u0026ldquo;the improvement headroom in DataFusion\u0026rsquo;s current physical optimizer on these queries\u0026rdquo;, not as \u0026ldquo;LLM beats cost-based optimizer.\u0026rdquo; That framing matters, because the rest of this post is about architectural choices, not about who wins on accuracy.\nWhat actually makes route C interesting relative to route B is three properties that have nothing to do with accuracy.\nArgument 1 · A patch is the smallest unit you can review Anyone who has written a Catalyst rewrite rule in Spark (PushDownPredicates, ColumnPruning, and friends) knows a common review difficulty: after a rule mutates a LogicalPlan, the diff is structural, and it is hard to see at a glance which fibre the rule actually pulled.\nJSON Patch gives you a smaller unit:\n[ { \u0026#34;op\u0026#34;: \u0026#34;replace\u0026#34;, \u0026#34;path\u0026#34;: \u0026#34;/nodes/3/build_side\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;left\u0026#34; }, { \u0026#34;op\u0026#34;: \u0026#34;remove\u0026#34;, \u0026#34;path\u0026#34;: \u0026#34;/nodes/5\u0026#34; } ] Each patch is atomic. It can be reverted independently. It can have a single regression test pinned to it. From an engineering-governance standpoint, I\u0026rsquo;d argue this is more controllable than \u0026ldquo;we wrote a new rule\u0026rdquo; or \u0026ldquo;the agent picked this join order\u0026rdquo; — particularly compared to route B, where the LLM agent emits \u0026ldquo;use this join order\u0026rdquo; with no diff. You either trust the agent\u0026rsquo;s choice or you don\u0026rsquo;t.\nThis is not an argument that patches are always correct. It is an argument that when they\u0026rsquo;re wrong, you can point at exactly which patch was wrong.\nArgument 2 · OLAP queries repeat — API cost amortises Per-query LLM cost is the first engineering objection to any \u0026ldquo;LLM in the optimizer\u0026rdquo; proposal.\nThe DBPlanBench paper reports per-query optimization cost typically in \u0026ldquo;a few cents\u0026rdquo;2. I can\u0026rsquo;t independently reproduce that number, but it points at a useful framing: OLAP workloads have one defining property — the same query template runs over and over. Dashboards, scheduled reports, ETL — running a query a thousand times is normal. If a few cents of LLM call buys a patch that\u0026rsquo;s reused a thousand times, the economics work.\nThe paper also has a detail that supports this view: an optimization the LLM found on SF=3 was carried over to SF=10 via a one-shot deterministic rewrite script2, retaining the speedup. The scalability of that mechanism still needs validation on bigger data (the paper doesn\u0026rsquo;t show SF=100/1000), but the pattern it points at is the right one: find once with an LLM, reuse as a rule for a long time. Not invoke the LLM on every execution.\nBy contrast, in route B the agent goes through its reasoning loop every time. The amortisation path is not natural.\nTo be honest: the paper does not include a \u0026ldquo;after how many reuses does the patch pay for itself\u0026rdquo; experiment either. Until someone publishes that curve on real workloads, \u0026ldquo;amortisable\u0026rdquo; is a hypothesis, not a conclusion.\nArgument 3 · Patch caches can become real infrastructure If you accept the \u0026ldquo;find once, reuse many\u0026rdquo; model, the patch is no longer a one-off optimization result. It is an asset — indexable by query signature, storable, auditable.\nThat opens several natural engineering extensions:\nDedupe by query signature: same query template with different literals, the same patch likely still applies. A/B and shadow execution: a patch can hang off the plan, run as shadow, only become active after measured wins. Audit trail: every applied patch is logged, so \u0026ldquo;which query was changed by whose patch\u0026rdquo; is answerable. This is well-aligned with practices SQL gateways already run — signature-keyed plan caches, plan-level audit logs. The LLM\u0026rsquo;s role here is closer to \u0026ldquo;patch generator\u0026rdquo; than \u0026ldquo;runtime decision maker.\u0026rdquo;\nWhere to plug in on Spark — the hook already exists If you wanted to actually try this on Spark, the API surface is already there. SparkSessionExtensions exposes injectPlannerStrategy, injectOptimizerRule, and injectPostHocResolutionRule extension points4; in particular, injectPlannerStrategy lets you register a SparkStrategy that runs during logical → physical conversion — exactly after the optimizer has done its work and before execution begins. A patch applier can plug in there.\nThe real engineering bottleneck is not the API. It is serialisation/deserialisation between SparkPlan and JSON. Catalyst plan tree nodes don\u0026rsquo;t ship with an official JSON codec.\nFrom what I can tell, that codec is the prerequisite the patch route needs before anything else lands. DBPlanBench\u0026rsquo;s flat node-id schema on DataFusion (node id + input/left/right references) is a reasonable reference point.\nLLM-PM is a separate route, not a cheaper C There is one obvious objection to address head-on: if you want an LLM helping pick plans, isn\u0026rsquo;t training-free plan retrieval cheaper?\nThere is such work. LLM-PM (arXiv:2506.05853)5 uses text-embedding-3-large to embed EXPLAIN plan text into vectors; for a new query it runs KNN against a history of past plans and applies the most similar historical plan\u0026rsquo;s shape. Fully training-free. On OpenGauss + JOB-CEB the paper reports mean −21% latency.\nLook at the distribution: about 20% of queries are slowed down, 60% unchanged5. The mean −21% is driven by a minority of queries getting a large speedup — which is a well-known reporting-discipline problem in this area. Any \u0026ldquo;mean +N%\u0026rdquo; result deserves the follow-up: \u0026ldquo;what\u0026rsquo;s the speed-up / slow-down / no-change distribution?\u0026rdquo;\nBut distribution isn\u0026rsquo;t the main point. The main point is that LLM-PM and DBPlanBench are not solving the same problem:\nLLM-PM = retrieve and apply an existing plan. It cannot create new structure. Whatever isn\u0026rsquo;t in the history library falls back to baseline. DBPlanBench = generate new plan variants. It can produce structures the original optimizer never explored, at the cost of paying a GPT-5 call. So this isn\u0026rsquo;t \u0026ldquo;cheaper vs. more expensive.\u0026rdquo; It\u0026rsquo;s two different jobs: retrieval reuses what\u0026rsquo;s known; generation fills what\u0026rsquo;s missing. A realistic deployment probably stacks them — try KNN first, fall back to LLM patch generation on miss — rather than choosing one.\nClosing — what the next experiment should actually measure The discussion of \u0026ldquo;LLM in the query optimizer\u0026rdquo; has often stalled at \u0026ldquo;can it replace the cost-based optimizer?\u0026rdquo; That may be the wrong question.\nPutting it behind the optimizer, doing last-mile tuning via patches, looks better than the alternatives on three engineering properties: small output granularity is review-friendly; the amortisation model is clean and fits OLAP\u0026rsquo;s repeating-query nature; and the landing hook in engines like Spark already exists (SparkSessionExtensions.injectPlannerStrategy).\nBut to move this from \u0026ldquo;engineering-plausible\u0026rdquo; to \u0026ldquo;benchmark-proven,\u0026rdquo; what\u0026rsquo;s missing isn\u0026rsquo;t more framing — it\u0026rsquo;s a concrete set of numbers. If I had to name what the next round of work should report, it would be these:\nSpark CBO + AQE as baseline, TPC-DS SF≥100, how much head-room is left for patches on top-cost queries. 4.78× is from DataFusion at SF=10; what\u0026rsquo;s left on a mature CBO at TB scale is the real question. Patch cache hit-rate curves: same template, different literals — how often the patch still applies; when it doesn\u0026rsquo;t, what selectivity drift caused the miss. Payback curves: how many reuses, on average, does a patch need to cover its LLM-generation cost, on real dashboard traffic. Regression rate: corresponding to LLM-PM\u0026rsquo;s 20% slowdown share, what does the patch route show on the same benchmarks, and is there auto-detection + auto-revert. Until those numbers exist, this post is a direction, not a conclusion. If I were running the experiment myself, the first thing I\u0026rsquo;d build is the SparkPlan ↔ JSON codec — the LLM, the rule-isation, the cache all attach to it, but without the codec the route can\u0026rsquo;t take a single step.\nDatabricks Blog. \u0026ldquo;Are LLM agents good at join order optimization?\u0026rdquo; 2026-04-22. https://www.databricks.com/blog/are-llm-agents-good-join-order-optimization\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nErol M.H., Hao X., Bianchi F., Greco C., Tagliabue J., Zou J. \u0026ldquo;Making Databases Faster with LLM Evolutionary Sampling.\u0026rdquo; arXiv:2602.10387. https://arxiv.org/abs/2602.10387 · Code (MIT): https://github.com/BauplanLabs/Making-Databases-Faster-with-LLM-Evolutionary-Sampling\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nIETF RFC 6902: JavaScript Object Notation (JSON) Patch. https://datatracker.ietf.org/doc/html/rfc6902\u0026#160;\u0026#x21a9;\u0026#xfe0e;\nApache Spark API: org.apache.spark.sql.SparkSessionExtensions — injectOptimizerRule / injectPlannerStrategy / injectPostHocResolutionRule. https://spark.apache.org/docs/latest/api/java/org/apache/spark/sql/SparkSessionExtensions.html\u0026#160;\u0026#x21a9;\u0026#xfe0e;\narXiv:2506.05853, \u0026ldquo;Training-Free Query Optimization via LLM-Based Plan Similarity.\u0026rdquo; https://arxiv.org/abs/2506.05853\u0026#160;\u0026#x21a9;\u0026#xfe0e;\u0026#160;\u0026#x21a9;\u0026#xfe0e;\n","permalink":"https://yaooqinn.github.io/posts/query-engines/llm-as-plan-tuner-not-optimizer/","summary":"Putting the LLM after the optimizer, emitting JSON patches for local plan tuning, is easier to reason about as engineering than asking it to replace the cost-based optimizer.","title":"LLMs Shouldn't Replace the Query Optimizer — They Should Sit Behind It"},{"content":"A paper landed on arXiv last week that is worth a second look from anyone who works on query engines: Finding Performance Issues in Database Systems by Exploiting Dormant Code Paths by Jinsheng Ba and Zhendong Su (ETH Zürich), arXiv:2605.22992, submitted 21 May 2026.\nThe headline finding is concrete: 21 previously-unknown, unique performance issues across PostgreSQL, MySQL, CockroachDB, and MariaDB, all surfaced with the workloads of TPC-H and TPC-DS. The technique that produced them — Branch Flip Analysis, or BFA — is conceptually simple, and the surface area in Apache Spark is unusually inviting. This post walks through what the paper does, why it works, and what an honest port to Spark would look like.\nWhat BFA actually is The methodology is one sentence: take every if branch that exists to turn an optimization on or off, flip it, and check whether the flipped path is significantly faster than the original path. If it is, you have a performance bug — the optimization, by definition, is supposed to be at least neutral.\nThe authors implement this in a prototype called QueryZen and apply it to four mature open-source DBMSs. Three details make the method usable in practice rather than a flood of false positives:\nDifferential testing for soundness. A flipped branch must keep query results unchanged. QueryZen runs the original and flipped binaries against a corpus and discards any flip that changes outputs. The authors also manually inspect a sample to argue that the branches that survive are genuinely optimization-only (logging, caching, plan choice) rather than functional. Statistical significance. A flip is reported only when the flipped runtime is significantly faster than the original across repeated trials. This filters out noise inherent to wall-clock benchmarking. Triage-friendly output. Each finding is a (commit, branch, query) tuple, which means a developer can reproduce it with a single configuration toggle rather than re-deriving the workload. Two limits the paper is honest about: results that look consistent on a given machine may shift on different hardware or memory configurations, and \u0026ldquo;true positive\u0026rdquo; is defined as \u0026ldquo;confirmed by the upstream developers\u0026rdquo; — a high bar that makes the 21 number a lower bound.\nWhy this is not the same as A/B-ing a config flag Database engines have always had configuration knobs that disable optimizations: enable_hashjoin = off in PostgreSQL, optimizer_switch in MySQL, the various spark.sql.* toggles in Spark. Practitioners have been flipping these for decades to debug regressions.\nWhat BFA adds is scale and coverage: it does not rely on the optimizations that happen to have a config flag exposed. It targets every if in the source tree that the authors can classify as an optimization branch — including ones that no user-facing knob controls. Most of the 21 issues live in branches that had no flag. That is the part of the result that should make engine maintainers pay attention: the well-known knobs have been A/B-tested to death, but the un-flagged optimization branches buried deep in the planner and executor have not.\nThe Spark surface area Here is where this gets interesting for the Spark community. Two ingredients BFA needs are already shipped in mainline Spark:\nA first-class way to disable a single optimizer rule. spark.sql.optimizer.excludedRules takes a comma-separated list of fully-qualified rule class names and removes them from the Catalyst optimizer batches. This is exactly the \u0026ldquo;branch flip\u0026rdquo; knob BFA wants — granular, declarative, and already maintained. A canonical benchmark workload. The TPC-DS benchmark generator in sql/core is already used in the project for performance regression work, and a public 99-query TPC-DS run is a familiar baseline in Spark performance discussions. So the recipe for a Spark-side BFA is not exotic:\nEnumerate the rules in Optimizer.batches and the rules registered via the SparkSessionExtensions API for a given build. For each rule, run a fixed TPC-H or TPC-DS workload twice: once with the rule enabled (the default), once with the rule listed in spark.sql.optimizer.excludedRules. Take the median of N trials. Verify that the produced query results are identical between the two runs. Discard the pair otherwise. Report rule × query pairs where the rule-disabled run is significantly faster than the rule-enabled run. There are honest complications. Optimizer rules in Catalyst are not all order-independent — disabling rule A may change whether rule B\u0026rsquo;s pattern matches, so the local \u0026ldquo;what happens if I flip just this one\u0026rdquo; semantics can leak. AQE makes some of these decisions at runtime based on observed statistics, which means a deterministic flip needs deterministic inputs. And the cost of running TPC-DS once per rule, times the number of rules in modern Catalyst, is not trivial. None of those make the experiment impossible; they make it engineering work.\nWhat this is not It is worth being precise about what BFA does and does not promise.\nIt is not a proof of optimality. Finding 21 bugs in four mature systems is impressive, but BFA cannot tell you the optimizer is correct everywhere it did not flag. It is not a substitute for cost-based reasoning. The flipped branch is faster on a particular query under a particular configuration. The \u0026ldquo;fix\u0026rdquo; might be a change to a cost model or a heuristic threshold, not a wholesale removal of the optimization. Two of the false-positive categories the paper discusses are exactly this — branches whose authors made deliberate trade-offs that BFA cannot read. It is not workload-free. TPC-H and TPC-DS bias toward analytics. A BFA pass on Spark would likely miss issues that only manifest under streaming, structured streaming, or DSv2 connector-heavy workloads. A modest proposal, framed as a question I am writing this in a personal capacity. But I think it is worth asking, in public, whether the open-source query engine community — Spark, Trino, DuckDB, Velox-backed engines, ClickHouse — should treat per-rule perf regression testing as part of the CI contract for new optimization work, the way most projects already treat correctness regression testing.\nThe materials are there. excludedRules already exists in Spark. Trino has optimizer.\u0026lt;rule\u0026gt;.disabled flags for many of its optimizers. DuckDB exposes SET disabled_optimizers. Most major projects already run TPC-H or TPC-DS nightly. The missing piece is the harness: enumerate, flip, diff, report.\nThe cost of doing this is real — both compute time and the engineering work of distinguishing genuine regressions from deliberate trade-offs. The cost of not doing it is the kind of thing BFA just demonstrated: 21 bugs sitting in mature, well-reviewed codebases, in branches that nobody had a reason to specifically look at.\nFurther reading The paper: arXiv:2605.22992 (PDF, HTML versions linked from the abs page). For Spark\u0026rsquo;s per-rule disable mechanism, see spark.sql.optimizer.excludedRules in the Spark 3.5+ configuration reference. For prior work on automated DBMS performance testing — APOLLO (Jung et al., 2019), AMOEBA (Liu et al., 2022), CERT (Ba et al., 2024) — the BFA paper\u0026rsquo;s related-work section is a compact map. These are personal observations from someone who works on Apache Spark, not a project position. The materials linked above are the primary sources; please form your own view from them.\n","permalink":"https://yaooqinn.github.io/posts/spark/branch-flip-analysis-from-postgres-to-spark/","summary":"An ETH paper finds 21 previously unknown performance bugs in PostgreSQL, MySQL, CockroachDB and MariaDB by flipping optimization branches on and off. The technique is conceptually simple, the surface in Spark is unusually inviting, and the open-source engine community already ships one of the building blocks.","title":"Branch Flip Analysis: A White-Box Way to Find Performance Bugs, and What It Means for Spark"},{"content":"My previous post covered the Databricks × UPenn LLM-for-join-order experiment, which argued that an LLM agent can do useful work on offline join-order tuning because it bypasses the two layers of approximation in a cost model — observing real execute(plan) wall-clock directly, trading trial-and-error budget against model bias.\nThe subtext of that post was: \u0026ldquo;give the model the right signals and enough feedback, and the LLM earns its keep inside a query engine.\u0026rdquo; This post is the mirror image — what happens when you don\u0026rsquo;t feed, don\u0026rsquo;t train, don\u0026rsquo;t give feedback, and just call a commercial LLM API to rewrite a SQL string.\nThe answer is a little embarrassing.\nOne embarrassing number I recently read E³-Rewrite (arXiv:2508.09023, August 2025), which targets LLM-driven SQL rewriting under three goals — Executable / Equivalent / Efficient. The contribution is a GRPO-based RL fine-tuning pipeline, which I\u0026rsquo;ll come back to. But the number that actually made me stop was the \u0026ldquo;baseline\u0026rdquo; row in their Table 1:\nMethod TPC-H 10GB avg latency (s) TPC-H p90 latency (s) Original (no rewrite) 78.81 300.00 LLM-only (GPT-4o) 74.92 300.00 LearnedRewrite (SIGMOD'22) 41.34 103.41 LLM-R² 54.76 300.00 R-Bot 39.89 84.27 E³-Rewrite (Qwen3-32B, post-training) 29.67 51.37 (Source: E³-Rewrite paper Table 1. PostgreSQL, TPC-H SF=10, ~2000 queries, 5 runs per query with min/max trimmed.)\nLook at the second row: directly asking what is widely considered the strongest commercial LLM (GPT-4o) to rewrite SQL via a prompt takes mean latency from 78.81s to 74.92s.\nEffectively no change. Statistically indistinguishable from \u0026ldquo;do nothing.\u0026rdquo;\nThe last row is the same paper\u0026rsquo;s own method: same PostgreSQL, same workload, swap the model for a 14B/32B open-source Qwen, feed it plan signals, train it once with RL, and the mean drops from 74.92 to 29.67 — down to 38% of the original.\nThe gap is not an LLM gap (GPT-4o is in fact stronger than Qwen3-32B on most general tasks). It\u0026rsquo;s a signal gap.\nWhy the LLM knows SQL but can\u0026rsquo;t rewrite SQL This looks counterintuitive on the surface: LLM training corpora are stuffed with SQL, and rewriting a query into an equivalent, faster form doesn\u0026rsquo;t sound that hard. But think it through.\nWhat does a human do before rewriting a SQL query?\nRun EXPLAIN to see the current plan Spot the full table scan, the nested-loop join Notice which filter chops 100M rows down to 1K Then decide: can this subquery be flattened? can this OR become a UNION ALL? can EXISTS become IN? The human gets more than the SQL string — they get the engine\u0026rsquo;s interpretation of that string: what\u0026rsquo;s slow, why, where the bottleneck is.\nStuffing only the SQL text into the prompt is like asking a human to optimize the query without looking at EXPLAIN. The LLM knows ten thousand rewrite patterns, but it has no idea which one matters for this query on this database right now. So it either guesses based on schema names and applies some generic rewrite, or it just echoes the SQL back unchanged.\n74.92 vs 78.81 — what that number is really telling you is: it either echoes the SQL or rewrites it into something that runs about the same; the fraction that actually wins is too small to show up in the mean.\nStep 1: feed it EXPLAIN The first thing E³-Rewrite does is exactly this — they call it Execution Hint Injection: before handing the SQL to the model, run EXPLAIN (at training time, EXPLAIN ANALYZE), take the plan, linearize the tree into indented text, and prepend it to the SQL.\nSo the prompt now contains something like:\nSeq Scan on lineitem (cost=0.00..172800 rows=6000000) Filter: l_shipdate \u0026lt; \u0026#39;1998-12-01\u0026#39;::date Rows Removed by Filter: 4500000 ← bottleneck -\u0026gt; Hash Join (cost=...) Hash Cond: (l_orderkey = o_orderkey) ... Now the model can see: the filter on lineitem is throwing away 75% of the rows, and it\u0026rsquo;s filtering after a full scan before the join. It now has a concrete optimization target — push the filter down, suggest an index, or rewrite the predicate.\nThe effect is immediate. From the same paper\u0026rsquo;s Table 4 ablation:\nE³-Rewrite variant avg latency (s) improved queries equivalence ratio Full 29.67 210 99.6% w/o execution hint 39.56 163 100% w/o RL 32.29 194 90.1% w/o demonstration retrieval 35.39 177 96.5% Vanilla (no fine-tune) 56.71 125 84.8% (Same setup, TPC-H 10GB.)\nRemoving execution hints alone takes latency from 29.67 to 39.56 — this single component accounts for a 25% regression. A model that was rewriting nothing useful starts knowing where to push, just because it saw a piece of EXPLAIN text.\nThe general lesson: the model doesn\u0026rsquo;t lack ability, it lacks visibility. Any serious LLM × query-engine work needs some representation of the plan in the prompt — not necessarily EXPLAIN text (JSON, protobuf, an operator graph all work), but something. SQL string alone is optimization with the lights off.\nStep 2: train it with a reward Plans aren\u0026rsquo;t enough. If you fed EXPLAIN to GPT-4o, would it just work? Better, but still far short.\nBack to Table 4, third row: \u0026ldquo;w/o RL\u0026rdquo; — keep the plan hint, keep the demonstrations, but skip RL fine-tuning, use the base model directly. Result: 32.29s and 90.1% equivalence.\nWhat does 90.1% mean? For every 100 queries rewritten, 10 of them come back not equivalent to the original — different results, or syntax error. I don\u0026rsquo;t need to spell out what that means in production.\nE³-Rewrite\u0026rsquo;s training objective is straightforward, three reward components:\nExecutability: does the rewritten SQL actually run (syntax + schema check) Equivalence: does it return the same result as the original Efficiency: given executable + equivalent, how much faster (cost-model estimate or actually run it) The three components are weighted, and GRPO normalizes within a group of candidate rewrites. The curriculum is restrained: stage 1 emphasizes executability + equivalence (force it to rewrite correctly first), stage 2 adds the efficiency reward — avoiding the trap of optimizing for speed at the expense of correctness from day one.\nAfter training, equivalence climbs from 90.1% to 99.6%, and latency drops from 32.29 to 29.67. This step isn\u0026rsquo;t icing — it\u0026rsquo;s the line between \u0026ldquo;usable\u0026rdquo; and \u0026ldquo;not usable.\u0026rdquo;\nStep 3: train on truth, infer on estimate This is an engineering detail in E³-Rewrite I think is underrated — the paper covers it in a single short paragraph, but it\u0026rsquo;s worth pulling out.\nEXPLAIN ANALYZE in PostgreSQL actually runs the query and gives you true per-operator row counts and timing. Plain EXPLAIN only gives the optimizer\u0026rsquo;s estimated cost, without running.\nE³-Rewrite\u0026rsquo;s setup: training time uses EXPLAIN ANALYZE for ground-truth plans (slower at training is fine, sample quality matters more); inference/deployment uses plain EXPLAIN (production has to be fast — you can\u0026rsquo;t run the original query before rewriting it).\nThis is a plain but generalizable pattern: give the model oracle signals during training, proxy signals during inference, and let it learn to infer the oracle from the proxy. It\u0026rsquo;s effectively a flavor of distillation, but it lands naturally on systems × LLM problems — almost every query optimization task has the same oracle/proxy asymmetry:\ncost model: real wall-clock during training, estimate at inference cardinality estimation: real row counts during training, stats lookup at inference index recommendation: what-if at training, no what-if at inference A pattern worth remembering.\nThree questions still open The thesis is in place — \u0026ldquo;LLM rewriting SQL\u0026rdquo; is bottlenecked by the signal path, not by the LLM. But the paper\u0026rsquo;s recipe is still a step or two away from landing in an open-source engine. I\u0026rsquo;ll leave three open questions:\n1. Is there a standard interface for feeding a plan to an LLM?\nEXPLAIN text is the most primitive option. It has versioning issues (different engines, different versions, different formats), information loss (much of the optimizer state never surfaces in EXPLAIN), and token waste (a deep plan linearized to text is long). If LLM × query engine is going to industrialize, this layer probably needs an engine-neutral plan representation — protobuf? Substrait? Some LLM-friendly intermediate? No consensus yet.\n2. There\u0026rsquo;s no consensus tooling for equivalence verification.\nE³-Rewrite\u0026rsquo;s reward function mentions \u0026ldquo;equivalence verification\u0026rdquo; but doesn\u0026rsquo;t say exactly how — is it SPES? SQLancer? LLM-as-judge? Each has a ceiling: SPES is precise but only over a SQL subset; SQLancer is strong but it\u0026rsquo;s differential testing — finds counterexamples, doesn\u0026rsquo;t prove equivalence; LLM judge is flexible but may hallucinate. Without a reliable equivalence verifier, no LLM rewrite proposal will pass a production review. This infrastructure gap is glaring.\n3. Will the conclusions on a 10GB benchmark hold — or flip — at TB scale and on real multi-tenant workloads?\nE³-Rewrite runs TPC-H SF=10 (10GB), which is on the larger side academically but still two to three orders of magnitude shy of production. Two things change with scale: rewrites yield bigger absolute wins (slow queries are slower; the optimization headroom is larger), and side effects also scale (shuffle volume, memory pressure shift). Whether a rewrite that wins -62% on small data lands as -80% or +10% on TB-scale data — there is no public data to answer this today.\nI don\u0026rsquo;t have an actionable takeaway to leave you with. The entire point of this post is to put the 74.92 number at the entrance of every \u0026ldquo;use LLM to rewrite SQL\u0026rdquo; project:\nIf your design doesn\u0026rsquo;t show the LLM the plan, and doesn\u0026rsquo;t use any form of feedback to train it or filter it, then it doesn\u0026rsquo;t matter whether you\u0026rsquo;re calling GPT-4o, Claude 4, or Gemini 3 — what you build will be statistically indistinguishable from doing nothing.\nWhether LLMs can rewrite SQL is a signal-path question, not a model-strength question.\n","permalink":"https://yaooqinn.github.io/posts/query-engines/llm-only-rewrite-doesnt-work/","summary":"On TPC-H 10GB, asking GPT-4o to rewrite SQL takes mean execution time from 78.81s down to 74.92s — almost nothing. Swap in an open 14B model, feed it plans, add a reward, fine-tune once, and the same workload drops to 29.67s. Whether LLMs can help SQL rewriting is not a question about model strength; it\u0026rsquo;s a question about whether you\u0026rsquo;re willing to give the model the signals it actually needs.","title":"Just Asking an LLM to Rewrite SQL Does Almost Nothing"},{"content":"In April 2026, Databricks Engineering published a piece with UPenn called Are LLM agents good at join order optimization?. The headline numbers were eye-catching: on the 113 queries of the Join Order Benchmark (JOB), the join orders chosen by an LLM agent dropped Databricks Runtime\u0026rsquo;s P90 latency by 41% and delivered a 1.288× geomean speedup — better than what you get by feeding the optimizer perfect cardinality estimates, and better than BayesQO, an offline Bayesian optimizer.\nIf you have spent any time around query engines, your first reaction is probably skepticism. After reading the original post and the UPenn follow-up, most of mine went away — but a few observations are worth pulling out, especially in the context of an open-source engine like Apache Spark.\nWhat the study actually does Stripped of marketing, the setup is surprisingly modest:\nThe agent does not enter the hot path. It does not replace the cost-based optimizer. It does not run inside the ms-level query compiler. It does what human DBAs have done by hand for decades — try plans offline, repeatedly, with feedback — and automates that loop. A single tool. The agent has exactly one: execute(plan), which returns wall-clock runtime plus the true cardinality of each subtree. Structured output. A grammar constrains the agent to emit only well-formed join orders, eliminating the need for any LLM-output validation logic. Adjustable budget. This is an anytime algorithm. The reported numbers use 15 rollouts per query; the agent keeps improving up to ~50 iterations. The dataset is IMDb scaled 10× to about 20GB. Baselines include the default DBR optimizer, the same optimizer fed perfect cardinality estimates, a smaller-model agent, and BayesQO. The result is the headline above.\nOne case study (q5b) tells the story well. It is a 5-way join. The default DBR optimizer starts from \u0026ldquo;American VHS company\u0026rdquo; (12 rows, looks the most selective). The agent flips it around and starts by filtering \u0026ldquo;VHS releases referencing 1994\u0026rdquo; — because a LIKE predicate caused the CBO to mis-estimate cardinality there by an order of magnitude.\nWhy \u0026ldquo;beating perfect cardinality\u0026rdquo; is not as mystical as it sounds This is the most easily-misread claim in the post.\n\u0026ldquo;Perfect cardinality estimates\u0026rdquo; sounds like a mathematical lower bound — if every step\u0026rsquo;s true row count is given to the optimizer, how can an LLM still win?\nThe answer: the cost model itself is not perfect.\nEvery cost-based optimizer relies on two approximations:\ncardinality → cost (CPU + I/O + network + shuffle) → wall-clock time ^approx 1 ^approx 2 The paper eliminates the first one by injecting perfect cardinalities. But the second one — the mapping from cost to actual time — is still there. The constants inside any cost model (shuffle rate, broadcast thresholds, spill penalties) reflect the hardware assumptions of the era they were tuned in. On NVMe / large memory / columnar / vectorized engines, those constants are routinely off.\nThe agent wins by bypassing both approximations: it observes the wall-clock returned by execute(plan) directly, trading trial-and-error cost for model bias.\nThat is not the LLM transcending mathematical optimality. It is trial-and-error beating modeling — a statement that only carries weight once you accept that the cost model is itself imperfect.\nA three-tier ladder for LLMs × query optimizers Flatten every \u0026ldquo;LLM + query optimization\u0026rdquo; proposal you have ever seen, and they tend to fit into one of three tiers:\nTier Role Latency budget Engineering difficulty Realistic today? L1 Hot-path replace Generate the final plan during compilation ms Very high (LLMs are too slow, expensive, non-deterministic) ❌ Not realistic yet L2 Hint generation (online) Provide hints (join order, broadcast choice) at compile time tens of ms ~ seconds High (must coexist with ms-level compiler) ⚠️ Partially feasible, needs to be async L3 Offline tuning + hint writeback Pick queries offline, tune with an agent, persist results as hints / MVs / stats minutes ~ hours Medium (fully offline, no SLA pressure) ✅ Now demonstrated Databricks picked L3. That is a deeply pragmatic engineering call, worth unpacking:\nIt does not block the query path. Any online LLM call has 100+ ms of TTFB, which alone destroys interactive OLAP latency. L3 sidesteps this entirely. Failures are recoverable. If the agent picks badly, discard the hint and run the query the old way. The downside of an L3 miss is one extra query execution. L1 and L2 failures pollute production traffic. The ROI is quantifiable. \u0026ldquo;This query runs 200 times a day; the agent\u0026rsquo;s one-time tuning saves 30% per run\u0026rdquo; is the kind of arithmetic that easily justifies the budget. Naturally anytime. Strict-SLA queries get 5 rollouts; reports / ETL get 50. The \u0026ldquo;spend an hour tuning to save 30 minutes a day\u0026rdquo; model works. L3 is the most realistic productization path for LLMs × QO in 2026–2027. L2 becomes interesting once inference prices fall another order of magnitude. L1 needs longer still.\nWhat this means for the Apache Spark ecosystem Spark\u0026rsquo;s join reordering goes through JoinReorderDP (the DPhyp algorithm) plus cost-based estimation, and the weak points are the same as in DBR: LIKE / range / string predicate cardinality, and stale hardware constants inside the cost model. The path Databricks walked is reproducible on OSS Spark, and open source actually has structural advantages here:\nStatistics are accessible. DESCRIBE EXTENDED, ANALYZE TABLE, column histograms — the true cardinality feedback the agent needs is already exposed by Spark APIs. Plan history is accessible. Spark UI, Spark History Server, SQL execution event logs — the raw material for both training the agent and feeding it candidate queries already exists. Hint machinery is mature. /*+ BROADCAST(t) */, /*+ MERGE(t1, t2) */, /*+ SHUFFLE_HASH */, /*+ JOIN_REORDER */ — the plan an agent picks can be expressed as a hint without touching Catalyst. If someone wanted to take L3 to OSS Spark, the shortest path looks roughly like:\nMine the SQL execution event log for slow queries — recurring, long wall-clock, wide join trees. Build the single tool — an endpoint that accepts a join order, runs EXPLAIN ANALYZE, and returns the true cardinality and runtime of each child. Run N rollouts (start at 15, give SLA-loose queries up to 50). Decompile the winning plan into a Spark hint and persist it through any layer that can intercept SQL after parsing but before the optimizer — a gateway, a JDBC proxy, a SQL view layer, or a catalog-level hint table. Keep monitoring. When statistics shift or table distributions drift, the tuning should fire again. Open source can do this better than Databricks does, because the community could co-design a cross-engine hint-catalog spec. A note on vectorized execution (Gluten / Velox / Photon / Comet): join reordering still happens in Spark Catalyst. Native engines consume the plan; they do not rewrite it. L3 is a Catalyst-layer story. Getting it right on classic Spark already captures most of the upside, well before any vectorized engine work is required.\nOpen questions the post leaves unanswered The more carefully you read the piece, the more numbers you want to leave a question mark next to:\nHow much does it cost in tokens? 50 rollouts × N tokens × frontier-model price might exceed the query time saved. This is the core question for L3 viability, and the blog does not address it head-on. Which frontier model? The piece only says \u0026ldquo;frontier model\u0026rdquo; — no GPT-5 / Claude / Gemini, no version. That means the result is sensitive to a moving target and hard to reproduce. The dataset is small. 20GB of IMDb does not stand in for TPC-DS at 1TB. The generalization story for longer queries, wider fact tables, and large dimension tables is open. Which DBR version? Spark 3.5? DBR 15? Without knowing the baseline runtime, you cannot separate \u0026ldquo;the agent\u0026rsquo;s contribution\u0026rdquo; from \u0026ldquo;regular DBR runtime improvements.\u0026rdquo; No code release. Until UPenn or someone else publishes a minimum implementation that runs on PostgreSQL or Spark, the community cannot reproduce this independently. A closing note The most interesting thing about L3 is that it turns decades of hand-tuned DBA craft into an automatable loop, for the first time.\nThis story does not need anyone to bet on \u0026ldquo;LLMs replacing the query optimizer\u0026rdquo; — a vision that may never arrive. It only needs an LLM to be marginally better than heuristics at one verifiable, recoverable task: try plans, submit the best one. That bar has now been cleared.\nFor open-source engine communities, the signal is clear. L3 is a low-risk, quantifiable direction with real benchmark backing. What is needed next is for someone to do the open-source reproduction on Spark / Trino / DuckDB and close at least one or two of those five open questions.\nIf you are working on this, I would love to compare notes. Find me at @yaooqinn.\nReferences\nDatabricks Engineering · Are LLM agents good at join order optimization? · 2026-04-22 UPenn Sky-ADRs · How do LLM agents think through SQL join orders? Leis et al. · How Good Are Query Optimizers, Really? · VLDB 2015 (the original JOB paper) BayesQO — an offline Bayesian optimizer used as a baseline in the study ","permalink":"https://yaooqinn.github.io/posts/spark/llm-for-join-order-an-apache-spark-perspective/","summary":"Databricks and UPenn put an LLM agent to work as an offline join-order tuner and got P90 latency down 41% / geomean 1.288× speedup on JOB\u0026rsquo;s 113 queries — beating even perfect cardinality estimates. From the trenches of an open-source query engine, here is what that result does and does not prove.","title":"LLMs for Join Order: An Apache Spark Perspective on the Three-Tier Ladder"},{"content":"This is Part 5 of the Spark SQL Metrics deep-dive series:\nPart 1: Metric types, complete reference, and what they mean Part 2: How metrics work internally, and how AQE uses them for runtime decisions Part 3: Extension APIs, UI rendering, and REST API Part 4: How Gluten extends the metrics system Part 5 (this post): Gluten metrics internals — node mapping, pipeline aggregation, MetricsUpdaterTree, aggregation sub-phases, and shuffle metrics In Part 4, we saw what Gluten\u0026rsquo;s metrics look like from the outside — the 60+ counters that appear in the Spark UI. Now we go deeper: how those numbers actually get from native Velox operators back to the JVM. If you\u0026rsquo;ve ever stared at a confusing Gluten metrics value and wondered where it came from — or if you\u0026rsquo;re contributing to Gluten and need to add new metrics — this post is for you.\nSubstrait Node ID → Velox Operator Mapping When Gluten converts a Spark plan to a Velox plan via Substrait, every operator in the plan receives a plan node ID. This ID is the bridge between the JVM world (where Spark lives) and the C++ world (where Velox executes). The C++ side needs to map these IDs back to metrics arrays so that each native metric value ends up attached to the right Spark operator.\nHow getOrderedNodeIds() Works The key function is getOrderedNodeIds(). It performs a post-order traversal of the Velox plan tree, building an orderedNodeIds_ vector. The traversal order is critical — it determines which index in the flat metrics arrays corresponds to which operator.\nSpark Plan → Substrait Plan → Velox PlanNode tree ↓ getOrderedNodeIds() (post-order) ↓ orderedNodeIds_[0] = leaf operator orderedNodeIds_[1] = next operator ... orderedNodeIds_[N] = root operator Why post-order? Because the JVM-side MetricsUpdaterTree also walks children before parents. Using the same traversal order on both sides ensures indices stay synchronized without needing an explicit lookup table.\nSpecial Case: Filter–Project Fusion Velox fuses FilterNode → ProjectNode into a single FilterProject operator for better performance. When this happens, the Filter node has no separate runtime metrics — it\u0026rsquo;s been absorbed into the fused operator. Gluten handles this by adding the Filter\u0026rsquo;s plan node ID to omittedNodeIds_ and emitting zeros for its metrics slot. The JVM side sees the zero-filled slot and knows to skip it.\nThis is important to understand when debugging: if you see a FilterExecTransformer with all-zero metrics, it doesn\u0026rsquo;t mean the filter wasn\u0026rsquo;t evaluated — it means Velox fused it with its adjacent Project.\nSpecial Case: Union Velox represents a Union as LocalPartitionNode + LocalExchangeNode + fake ProjectNodes. This internal representation doesn\u0026rsquo;t map cleanly to a single Spark UnionExec. Gluten unwraps this structure to find the real children, ensuring the metrics arrays line up with the Spark plan\u0026rsquo;s structure rather than Velox\u0026rsquo;s internal representation.\nVelox Pipeline Model and Metrics Aggregation Here\u0026rsquo;s a subtlety that catches many developers off guard: Velox doesn\u0026rsquo;t execute a plan as a single pipeline. It splits the plan into multiple pipelines at exchange boundaries (and sometimes at other points like hash join build sides). A single logical operator can have instances running in different pipelines, and each instance collects its own metrics.\nHow toPlanStats() Aggregates toPlanStats(taskStats) gathers metrics from all pipeline instances and returns a Map[PlanNodeId → PlanStats]. Each PlanStats contains:\noperatorStats: a Map[SequenceId → OperatorStats], where each entry represents one pipeline instance of the operator When Gluten\u0026rsquo;s collectMetrics() iterates through these entries, it writes each pipeline instance into a separate metrics index:\nfor (const auto\u0026amp; entry : stats.operatorStats) { // Each entry is one pipeline instance of this operator metrics_-\u0026gt;get(Metrics::kWallNanos)[metricIndex] = entry.second-\u0026gt;cpuWallTiming.wallNanos; metricIndex++; } This means a single Spark operator might map to multiple metrics indices in the arrays. For example, a HashAggregateExec that appears in both a pre-shuffle partial aggregation pipeline and a post-shuffle final aggregation pipeline will have two separate metrics entries.\nJVM-Side Merging The JVM-side MetricsUpdater handles multi-pipeline entries by fetching all entries from the relMap for a given operator and calling mergeMetrics(). For timing metrics, this typically sums them. For peak memory, it takes the maximum. The merged result is what you see in the Spark UI — a single set of numbers that represents the operator\u0026rsquo;s total work across all pipelines.\nMetricsUpdaterTree Walking On the JVM side, MetricsUtil.scala orchestrates the entire metrics dispatch with two key methods.\nBuilding the Tree: treeifyMetricsUpdaters() treeifyMetricsUpdaters(plan) builds a MetricsUpdaterTree from the SparkPlan. This isn\u0026rsquo;t a simple recursive copy — there are several adjustments:\nHashJoin handling: The tree separates build and stream children, since Velox executes them in different pipelines with different metrics SortMergeJoin handling: Similarly separates buffer and stream children MetricsUpdater.None operators: These are skipped entirely — their child is linked directly to their parent. This happens for operators that Gluten replaces with no-ops (e.g., certain adapter nodes) Children are reversed: This is crucial. The children list is reversed to match the post-order traversal used by getOrderedNodeIds() on the C++ side Walking the Tree: updateTransformerMetricsInternal() updateTransformerMetricsInternal() walks the MetricsUpdaterTree and dispatches metrics to type-specific updaters:\nUpdater Operator Special Handling HashAggregateMetricsUpdater HashAggregate 3-phase sub-metrics (see next section) JoinMetricsUpdaterBase HashJoin Extra metrics entry for build phase SortMergeJoinMetricsUpdater SortMergeJoin Buffer/stream phase separation LimitMetricsUpdater Limit over Sort Skips Limit\u0026rsquo;s own metrics (Velox TopN handles both) Default Everything else mergeMetrics() → updateNativeMetrics() For joins, there\u0026rsquo;s an important detail: Velox reports build-phase metrics as an extra entry beyond what the relMap directly provides. The join updater knows to extract this extra entry and attach it to the build-side metrics, which is why you\u0026rsquo;ll sometimes see accurate build-side timing even though the build and probe happen in different pipelines.\nFor Limit over Sort, Gluten skips the Limit\u0026rsquo;s own metrics entirely. Velox implements this as a TopN operator that handles both the sorting and the limit in one fused operation, so there\u0026rsquo;s only one set of metrics to report.\nAfter dispatching, the walker recursively processes children with updated operator and metrics indices, ensuring each child picks up from where the parent left off in the flat metrics arrays.\nAggregation Sub-Phase Metrics Hash aggregation in Velox is more nuanced than in vanilla Spark. It can execute in up to three phases, controlled by AggregationParams. Understanding this split is essential for diagnosing aggregation performance.\nPhase 1: Extraction (extractionNeeded = true) Pre-aggregation column extraction — for example, extracting fields from nested structs before they can be grouped and aggregated.\nMetrics:\nextractionCpuCount — CPU time for extraction extractionWallNanos — Wall clock time for extraction If extraction time is high relative to total aggregation time, your schema might benefit from flattening nested columns before aggregation.\nPhase 2: Aggregation (always present) The main hash aggregation work — hashing group keys, looking up or creating groups, and accumulating values.\nMetrics:\naggOutputRows — Number of output rows (i.e., distinct groups) aggWallNanos — Wall clock time for aggregation aggPeakMemoryBytes — Peak memory used by the hash table aggSpilledBytes — Bytes spilled when memory pressure triggers spilling flushRowCount — Intermediate rows flushed when the hash table gets too large loadedToValueHook — Pushdown aggregation count (an optimization where aggregation is pushed into the scan operator) flushRowCount is particularly useful for debugging: a high flush count means the hash table keeps exceeding its memory budget, causing intermediate results to be flushed and re-aggregated. This leads to extra work and slower queries.\nPhase 3: Row Construction (rowConstructionNeeded = true) Post-aggregation row assembly — for example, constructing output struct columns from the aggregated results.\nMetrics:\nrowConstructionCpuCount — CPU time for row construction rowConstructionWallNanos — Wall clock time for row construction How the Phases Map to Metrics Entries The updater walks the aggregationMetrics list in order, consuming one entry per phase:\naggregationMetrics[0] → extraction phase (if needed) aggregationMetrics[1] → aggregation phase aggregationMetrics[2] → row construction phase (if needed) This three-phase split is unique to Gluten/Velox — vanilla Spark\u0026rsquo;s HashAggregateExec reports a single aggTime that lumps everything together. With Gluten, you can pinpoint where in the aggregation pipeline the time is being spent.\nShuffle Metrics Gluten\u0026rsquo;s columnar shuffle has its own metrics layer, and the available metrics vary by shuffle writer type. Understanding which writer is in use tells you which metrics to look at — and which tuning knobs are relevant.\nBase Metrics (All Writers) Metric Display Name What It Measures dataSize data size Total shuffle data size bytesSpilled shuffle bytes spilled Bytes spilled during shuffle spillTime time to spill Time spent spilling compressTime time to compress Compression time decompressTime time to decompress Decompression time deserializeTime time to deserialize Deserialization time shuffleWallTime shuffle wall time Total shuffle wall clock time peakBytes peak bytes allocated Peak memory for shuffle buffers Hash Shuffle Writer Adds Metric Display Name What It Measures splitTime time to split Time splitting rows into partitions dictionarySize dictionary size Size of dictionary-encoded columns Sort Shuffle Writer Adds Metric Display Name What It Measures sortTime time to shuffle sort Time sorting rows by partition c2rTime time to shuffle c2r Time converting columnar→row format for sorting RSS (Remote Shuffle Service) Writer Adds Metric Display Name What It Measures sortTime time to shuffle sort Time sorting rows by partition Diagnosing Shuffle Bottlenecks The c2rTime metric deserves special attention. It represents the overhead of converting columnar batches to row format inside the sort-based shuffle writer. In columnar engines like Velox, data naturally lives in columnar format — converting it to rows for sorting is pure overhead.\nIf c2rTime dominates shuffleWallTime, the columnar-to-row conversion is your bottleneck. In this case, switching to hash-based shuffle (which can operate directly on columnar batches) might yield a significant speedup. This is one of the key decisions Gluten users face: hash shuffle is faster for wide tables with many columns, while sort shuffle uses less memory for high-cardinality partition keys.\nWrapping Up Gluten\u0026rsquo;s metrics machinery is complex because it bridges two very different execution models — Spark\u0026rsquo;s row-at-a-time Volcano iterator model and Velox\u0026rsquo;s pipeline-parallel vectorized model. The key concepts to remember:\nNode ID mapping via post-order traversal keeps the C++ and JVM sides synchronized Multi-pipeline aggregation means a single Spark operator\u0026rsquo;s metrics may come from multiple Velox pipeline instances MetricsUpdaterTree dispatches metrics to type-specific updaters that understand each operator\u0026rsquo;s internal structure Aggregation sub-phases give you visibility into extraction, aggregation, and row construction separately Shuffle writer type determines which metrics are available and which tuning strategies apply In Part 1, we covered the five metric types and the complete reference. In Part 2, we traced the internal lifecycle and AQE\u0026rsquo;s use of shuffle statistics. In Part 3, we explored extension APIs, UI rendering, and the REST API. In Part 4, we examined how Gluten extends the metrics system. In Part 5 (this post), we went deep into the internals — from Substrait-to-Velox node mapping to pipeline aggregation, MetricsUpdaterTree walking, aggregation sub-phases, and shuffle metrics. This concludes the series.\n","permalink":"https://yaooqinn.github.io/posts/spark/sql-metrics-part5-gluten-internals/","summary":"Part 5 of the SQL Metrics deep dive. How Gluten maps Substrait plan nodes to Velox operators, aggregates metrics across pipelines, walks the MetricsUpdaterTree, and handles aggregation sub-phases and shuffle metrics.","title":"Deep Dive into Spark SQL Metrics (Part 5): Gluten Metrics Internals"},{"content":"This is Part 6 of the Spark SQL Metrics deep-dive series:\nPart 1: Metric types, complete reference, and what they mean Part 2: How metrics work internally, and how AQE uses them for runtime decisions Part 3: Extension APIs, UI rendering, and REST API Part 4: How Gluten extends the metrics system Part 5: Gluten metrics internals — node mapping, pipeline aggregation, MetricsUpdaterTree Part 6 (this post): Metrics in action — a real-world TPC-DS q99 walkthrough with Gluten/Velox In Parts 1–5, we built up all the machinery: metric types, internal plumbing, extension APIs, Gluten\u0026rsquo;s architecture, and native-side aggregation. Now it\u0026rsquo;s time to put it all to use. We\u0026rsquo;ll open the Spark UI for a real query, walk through every operator\u0026rsquo;s metrics, and show exactly how to read them to understand what happened during execution.\nThe Query: TPC-DS q99 TPC-DS q99 is a 5-table join query that analyzes catalog sales shipping delays by warehouse, ship mode, and call center. It groups shipments into buckets based on how late they arrived (31–60 days, 61–90 days, 91–120 days, and over 120 days), then ranks the results.\nWe ran this query at scale factor 10,000 (10 TB of raw data) on a cluster using Gluten/Velox as the native execution backend, with Delta Lake tables on cloud object storage.\nThe query plan follows a classic star-schema pattern:\ncatalog_sales (fact table, 3.3B rows) → BroadcastHashJoin with date_dim → BroadcastHashJoin with ship_mode → BroadcastHashJoin with call_center → BroadcastHashJoin with warehouse → Partial HashAggregate → Shuffle (hash partitioning) → AQE Coalesce → Final HashAggregate → TakeOrderedAndProject (top 100) Every number in this post is real. Let\u0026rsquo;s walk through the metrics, operator by operator, and see what they tell us.\nSection 1: The Fact Table Scan — 3.3 Billion Rows The first operator in the plan is ScanTransformer catalog_sales. This is where everything starts, and the metrics here tell a rich story.\nScale number of raw input rows: 3,321,160,461 — 3.3 billion rows in the catalog_sales table number of output rows: 2,837,474,310 — 2.8 billion rows after predicate pushdown eliminates non-matching row groups size of files read: 910.9 GiB across 2,605 files spanning 1,837 partitions That\u0026rsquo;s a massive scan. But look at how efficiently it was served.\nThe I/O Tier — A Cache Story This is where it gets interesting:\nstorage read bytes: 0.0 B — zero bytes from remote storage local ssd read bytes: 2.1 GiB — all data served from local SSD cache ram read bytes: 0.0 B — no in-memory cache hits number of cache read bytes: 2.1 GiB — matches the local SSD number exactly Think about what this means: the table occupies 910.9 GiB on disk, but we only needed to read 2.1 GiB of actual data. That\u0026rsquo;s a 99.8% reduction. Two things made this possible:\nColumnar format efficiency — we only read the columns referenced by the query (a fraction of the table\u0026rsquo;s total columns) Predicate pushdown to row groups — Parquet/Delta statistics let Velox skip entire row groups whose min/max ranges don\u0026rsquo;t match the filter predicates And all 2.1 GiB came from local SSD cache — not a single byte traversed the network to remote storage.\nRow Group Pruning number of skipped row groups: 2,757 Velox examined row group statistics and skipped 2,757 row groups entirely. These row groups contained data outside the predicate ranges (the date_dim filter constrains d_month_seq to a specific range, which translates to a range on cs_sold_date_sk). This is the primary reason we read only 2.1 GiB from 910.9 GiB of files.\nDynamic Filters number of dynamic filters accepted: 9,380 This is a powerful optimization. During execution, the build side of each broadcast hash join produces a filter (a set of join key values or a Bloom filter). These filters are pushed back down to the scan operator at runtime. With 9,380 dynamic filters applied, the scan eliminated rows that would never match any join — before the rows even reached the join operators.\nIf you\u0026rsquo;ve read Part 4, you know Gluten instruments dynamic filter acceptance at the scan level. This metric tells you: yes, runtime filters were generated and applied, and they helped.\nTiming time of scan and filter: 4.5 minutes total across all tasks time of scan IO: 2.1 minutes — about half the scan time was I/O wait cpu wall time count: 969,724 batches processed The scan processed nearly a million batches. The I/O time (2.1 min) versus total scan time (4.5 min) tells us that roughly half the time was I/O and half was CPU work (decompression, predicate evaluation, column extraction).\nTask Distribution Looking at the min/median/max breakdown:\nBytes read per task: median 32 KiB, max 17.3 MiB — moderate skew in data distribution Peak memory per task: median 32 KiB, max 29.9 MiB Most tasks process small amounts of data (the table is split across 1,837 partitions), but some partitions are significantly larger. This is typical for real-world data — perfectly uniform partitioning is rare.\nSection 2: The Four Broadcast Hash Joins — 2.8 Billion Probe Rows After the scan, the plan chains four broadcast hash joins. Each joins the fact table with a dimension table: date_dim, ship_mode, call_center, and warehouse. Let\u0026rsquo;s walk through the date_dim join (node [18]) in detail, then summarize the pattern across all four.\nBuild Side (date_dim) number of hash build input rows: 855,925 (date_dim rows replicated across partitions via broadcast) time of hash build: 819 ms total — fast, small dimension table time of building hash table: 533 ms — actual hash table construction hash build peak memory bytes: 155.7 MiB total (68 KiB per task — tiny) The build side is trivially small. date_dim has a few hundred thousand rows, and after broadcast, each task gets a copy. Building the hash table takes well under a second.\nProbe Side Now the probe side — this is where the work happens:\nnumber of hash probe input rows: 2,837,474,310 — the entire 2.8 billion row output from the scan time of hash probe: 31.4 seconds total time of probing hash table: 4.1 seconds — actual hash lookups time of preparing hash table probe: 3.7 seconds — deserializing the broadcast data time of converting rows to columns: 8.3 seconds — columnar format conversion Notice the breakdown: out of 31.4 seconds of total probe time, only 4.1 seconds was spent on actual hash lookups. The rest was overhead — deserialization, format conversion, and pipeline coordination. This is typical for broadcast joins: the hash lookup itself is fast (the hash table fits in L2/L3 cache), but marshaling 2.8 billion rows through the pipeline takes time.\nDynamic Filters Generated number of hash probe dynamic filters produced: 2,345 Each task that processes this join produces a dynamic filter from the build side\u0026rsquo;s key values. These 2,345 filters are pushed back to the scan operator (contributing to the 9,380 dynamic filters we saw earlier — multiple joins contribute filters).\nNo Spill bytes written for spilling of hash build: 0.0 B bytes written for spilling of hash probe: 0.0 B Perfect. The dimension tables are small enough to fit entirely in memory. No data was spilled to disk during any join.\nOutput number of hash probe output rows: 2,837,474,310 — all 2.8 billion rows pass through Output data volume: 65.7 GiB (the output grows because we\u0026rsquo;ve added columns from the dimension table) All rows pass because the predicate pushdown and dynamic filters already eliminated non-matching rows during the scan. The join itself is just enriching each row with dimension attributes.\nPattern Across All Four Joins Here\u0026rsquo;s a summary of all four broadcast hash joins side by side:\nJoin Build Table Build Rows Probe Time Post-Projection Time date_dim date_dim 855,925 31.4s 8.1s ship_mode ship_mode 46,900 1.6 min 8.0s call_center call_center 126,630 2.4 min 8.1s warehouse warehouse 58,625 2.8 min 8.2s Notice how the probe time increases with each successive join. This isn\u0026rsquo;t because later joins are slower — it\u0026rsquo;s because wallNanos (which backs the probe time metric) includes child operator wait time. As we explained in Part 4, each operator\u0026rsquo;s wall time includes the time spent waiting for data from its child operators. The warehouse join (the outermost) includes the time of all three joins below it, plus the scan.\nThe post-projection time is remarkably consistent (~8 seconds each), which makes sense — each join appends a few columns, and the projection work is proportional to output size, which is roughly the same (2.8B rows) at each stage.\nSection 3: The Aggregation — 2.8B → 2.3M Rows After the four joins, the plan applies a FlushableHashAggregateExecTransformer (node [10]) for partial aggregation. This is where the data volume drops dramatically.\nnumber of output rows: 2,369,250 — a 1,200× reduction from 2.8 billion input rows time of aggregation: 6.6 minutes time of aggregate functions: 52.9 seconds — the actual SUM() computations time of preparing hash table probe: 5.4 minutes — dominated by child operator wait time (the joins above) peak memory bytes: 3.8 GiB total (max 3.6 MiB per task) number of spilled bytes: 0.0 B — no spill needed number of output vectors: 585 — only 585 output batches from 2.8 billion input rows The 1,200× reduction tells us that the group-by keys (warehouse, ship_mode, call_center, date bucket) have high cardinality but still produce far fewer groups than input rows. The actual aggregate computation (SUM) took only 52.9 seconds — the bulk of the 6.6-minute aggregation time is child wait time cascading up from the scan and joins below.\nWholeStageCodegenTransformer The entire native pipeline — from scan through all four joins through partial aggregation — runs inside a single WholeStageCodegenTransformer:\nduration: 26.2 minutes total (max 6.6 seconds per task) This is the end-to-end time for the native execution pipeline. It encompasses everything we\u0026rsquo;ve discussed so far: scanning 3.3 billion rows, four broadcast hash joins on 2.8 billion rows, and partial aggregation down to 2.3 million rows — all executed in Velox\u0026rsquo;s vectorized native engine.\nSection 4: The Shuffle — Hash-Based, Compact After partial aggregation, the plan shuffles data via ColumnarExchange (node [7]) to redistribute rows by their group-by keys for final aggregation.\nWrite Side shuffle bytes written: 243.6 MiB — remarkably small shuffle write time: 327 ms time to split: 33.7 seconds — hash partitioning into 512 partitions shuffle wall time: 21.7 seconds shuffle bytes spilled: 0.0 B — everything fits in memory peak bytes allocated: 255.1 GiB total (446.5 MiB max per task) The aggregation reduced 2.8 billion rows to 2.3 million rows, and those 2.3 million rows serialize to just 243.6 MiB. Compare that to the 910.9 GiB fact table we started with — the aggregation made the shuffle almost trivially small.\nThe split time (33.7s) is the time to hash-partition the output into 512 buckets. The actual write time (327 ms) is fast because there\u0026rsquo;s so little data to write.\nRead Side remote bytes read: 228.3 MiB local bytes read: 15.3 MiB remote reqs duration: 41.9 seconds — cross-node shuffle fetch time to deserialize: 2.2 seconds Most of the shuffle data (228.3 MiB) was read from remote executors, with a small fraction (15.3 MiB) from local tasks. The 41.9 seconds of remote request duration includes network latency and scheduling overhead — this is typical for cross-node shuffle on a distributed cluster.\nThe shuffle writer type is hash (visible in the plan), meaning rows are hash-partitioned by the group-by keys for the final aggregation.\nSection 5: AQE in Action — 512 → 149 Partitions After the shuffle, AQEShuffleRead (node [6]) kicks in:\nnumber of coalesced partitions: 149 — AQE merged 512 partitions down to 149 partition data size: 252.8 MiB total, target ~1.7 MiB per partition AQE examined the shuffle output statistics and determined that 512 partitions was too many for 252.8 MiB of data. At 512 partitions, each partition would average under 500 KiB — not worth the scheduling overhead of 512 tasks. By coalescing to 149 partitions, each task processes a more reasonable ~1.7 MiB.\nNo skewed partitions were detected (numSkewedPartitions is absent), so AQE only applied coalescing, not skew handling.\nThis is a perfect example of Part 2\u0026rsquo;s discussion of how AQE uses metrics for runtime decisions. The shuffle write statistics from Stage 1 directly informed Stage 2\u0026rsquo;s partition count.\nSection 6: Final Aggregation and Result — 4,050 → 100 Rows Final Aggregation RegularHashAggregateExecTransformer (node [4]) performs the final merge of partial aggregates:\nInput: 2,369,250 rows across 149 tasks number of output rows: 4,050 — the final group count time of aggregation: 3.2 seconds peak memory bytes: 194.4 MiB (1.3 MiB per task) No spill The partial aggregation already did the heavy lifting — the final aggregation just merges pre-aggregated groups. 2.3 million partial rows collapse to 4,050 final groups in 3.2 seconds. Trivial.\nTop-100 Sort and Result Delivery TakeOrderedAndProjectExecTransformer takes the 4,050 rows, sorts them, and returns the top 100. Then VeloxColumnarToRow converts the columnar result to row format for the driver:\nnumber of output rows: 100 From 3.3 billion rows to 100. That\u0026rsquo;s the entire query.\nThe Full Picture — What the Metrics Tell Us Let\u0026rsquo;s step back and see the complete execution flow with key metrics:\nScan catalog_sales: 3.3B raw → 2.8B rows (4.5 min, 2.1 GiB from SSD cache) ↓ 9,380 dynamic filters applied, 2,757 row groups skipped Join × 4 (broadcast): 2.8B rows through each join, zero spill ↓ Partial Aggregation: 2.8B → 2.3M rows (1,200× reduction) ↓ WholeStageCodegenTransformer: 26.2 min total native pipeline time Shuffle: 243.6 MiB written (hash, 512 partitions) ↓ AQE: 512 → 149 partitions (coalesced) ↓ Final Aggregation: 2.3M → 4,050 rows (3.2s) ↓ Top-100 Sort → 100 rows returned Key Takeaways from the Metrics 1. Cache is king. Zero bytes from remote storage, everything from local SSD. The 910.9 GiB table only needed 2.1 GiB of actual reads — a 99.8% reduction through columnar format efficiency and predicate pushdown.\n2. Row group pruning is effective. 2,757 row groups skipped. Velox used Parquet/Delta min-max statistics to eliminate entire row groups before reading a single byte from them. This is the primary driver of the I/O reduction.\n3. Dynamic filters work. 9,380 filters applied during the scan, generated at runtime from join build sides. These filters eliminated rows before they even reached the join operators, avoiding unnecessary processing of billions of rows.\n4. No spill anywhere. Joins, aggregations, and shuffle all fit in memory. Zero bytes spilled to disk across the entire query. The dimension tables were small enough for broadcast, and the partial aggregation reduced data volume before the shuffle.\n5. AQE coalescing helps. 512 → 149 partitions for the final aggregation. Without AQE, we\u0026rsquo;d have 512 tasks each processing less than 500 KiB — wasteful scheduling overhead for a trivial amount of data.\n6. The bottleneck is the native pipeline. 26.2 minutes in WholeStageCodegenTransformer, encompassing the scan of 3.3 billion rows, four joins on 2.8 billion rows, and partial aggregation. This is where the real work happens, and it\u0026rsquo;s all executed in Velox\u0026rsquo;s vectorized native engine.\n7. wallNanos increases up the tree. Each parent operator\u0026rsquo;s wall time includes its children\u0026rsquo;s wait time. The outermost join shows 2.8 minutes not because it\u0026rsquo;s slow, but because it includes the scan, three inner joins, and all their I/O. This confirms the caveat from Part 4 — always read wall time metrics with the operator tree in mind.\nWrapping Up This walkthrough demonstrates that metrics aren\u0026rsquo;t just numbers on a screen — they\u0026rsquo;re a narrative. Every metric answers a specific question:\nWhere did the data come from? → I/O tier metrics (all from SSD cache) How much work was avoided? → Row group pruning and dynamic filters Where did the data volume drop? → Aggregation (1,200× reduction) Was there any resource pressure? → Spill metrics (zero everywhere) Did the optimizer help? → AQE coalescing (512 → 149 partitions) When you next open the Spark UI for a slow query, you now have the vocabulary to read every metric, understand what it means, and pinpoint exactly where the bottleneck is.\nIn Part 1, we learned the five metric types and built a complete reference. In Part 2, we traced how metrics flow through the Spark internals and how AQE uses them. In Part 3, we explored the extension APIs, UI rendering, and REST endpoints. In Part 4, we saw how Gluten extends the metrics system for native execution. In Part 5, we went deep into Gluten\u0026rsquo;s native-side metrics machinery. And in Part 6 (this post), we put it all together — walking through a real TPC-DS query to show how every metric tells part of the execution story. This concludes the series.\n","permalink":"https://yaooqinn.github.io/posts/spark/sql-metrics-part6-in-action/","summary":"Part 6 of the SQL Metrics series. A real-world walkthrough of TPC-DS q99 at SF10000 with Gluten/Velox, reading every metric to understand what happened during execution.","title":"Deep Dive into Spark SQL Metrics (Part 6): Metrics In Action — TPC-DS q99 with Gluten"},{"content":"This is Part 1 of a 3-part series on Spark SQL Metrics:\nPart 1 (this post): Metric types, complete reference, and what they mean Part 2: How metrics work internally, and how AQE uses them for runtime decisions Part 3: Extension APIs, UI rendering, and REST API Part 4: How Gluten extends the metrics system What Are SQL Metrics? Every physical operator in Spark SQL can define metrics — counters that track what happened during query execution. When you click on a query in the Spark SQL tab and see numbers like \u0026ldquo;number of output rows: 5,000\u0026rdquo; or \u0026ldquo;peak memory: 512.0 MiB\u0026rdquo;, those are SQL metrics.\nThey are built on Spark\u0026rsquo;s AccumulatorV2 framework: each task updates its local copy, and the driver aggregates them after task completion.\nThe Five Metric Types Spark defines five metric types, each with different aggregation and display semantics:\n1. Sum (createMetric) The simplest type. Values from all tasks are summed into a single total.\nDisplay format: 1,234,567\nTypical usage: Row counts, file counts, partition counts.\n\u0026#34;numOutputRows\u0026#34; -\u0026gt; SQLMetrics.createMetric(sparkContext, \u0026#34;number of output rows\u0026#34;) 2. Size (createSizeMetric) For byte-based measurements. Shows the total plus per-task distribution.\nDisplay format: total (min, med, max): 512.0 MiB (128.0 MiB, 128.0 MiB, 128.0 MiB)\nTypical usage: Peak memory, spill size, data size, shuffle bytes.\n\u0026#34;peakMemory\u0026#34; -\u0026gt; SQLMetrics.createSizeMetric(sparkContext, \u0026#34;peak memory\u0026#34;) The (min, med, max) breakdown reveals per-task distribution — essential for detecting skew. If max is 10x the median, one task is doing most of the work.\n3. Timing (createTimingMetric) For millisecond durations. Shows total plus per-task distribution.\nDisplay format: total (min, med, max): 5.0 s (100 ms, 1.2 s, 2.0 s)\nTypical usage: Aggregation time, sort time, broadcast time, hash map build time.\n\u0026#34;aggTime\u0026#34; -\u0026gt; SQLMetrics.createTimingMetric(sparkContext, \u0026#34;time in aggregation build\u0026#34;) 4. Nanosecond Timing (createNanoTimingMetric) Same as timing but accepts nanosecond values, converted to milliseconds for display.\nDisplay format: Same as timing.\nTypical usage: Shuffle write time (measured in nanoseconds for precision).\n\u0026#34;shuffleWriteTime\u0026#34; -\u0026gt; SQLMetrics.createNanoTimingMetric(sc, \u0026#34;shuffle write time\u0026#34;) 5. Average (createAverageMetric) For per-task averages. Shows distribution of the average values across tasks.\nDisplay format: avg (min, med, max): (1.2, 2.5, 6.3)\nTypical usage: Hash probe efficiency.\n\u0026#34;avgHashProbe\u0026#34; -\u0026gt; SQLMetrics.createAverageMetric(sparkContext, \u0026#34;avg hash probes per key\u0026#34;) Reading the \u0026ldquo;total (min, med, max)\u0026rdquo; Format This is the most important format to understand:\npeak memory total (min, med, max) 512.0 MiB (128.0 MiB, 128.0 MiB, 128.0 MiB (stage 3.0: task 36)) Field Meaning total Sum across all tasks min Smallest task value med Median (50th percentile) max Largest task value, annotated with (stage X: task Y) Balanced workload: min ≈ med ≈ max\nSkewed workload: max \u0026raquo; med — investigate the annotated task\nComplete SQL Metrics Reference Scan Operators Metric Display Name Type Operators numOutputRows number of output rows sum DataSourceScanExec, DataSourceV2ScanExecBase, InMemoryTableScanExec, LocalTableScanExec numFiles number of files read sum DataSourceScanExec filesSize size of files read size DataSourceScanExec numPartitions number of partitions read sum DataSourceScanExec staticFilesNum static number of files read sum DataSourceScanExec staticFilesSize static size of files read size DataSourceScanExec metadataTime metadata time timing DataSourceScanExec scanTime scan time timing DataSourceScanExec pruningTime dynamic partition pruning time timing DataSourceScanExec Aggregation Operators Metric Display Name Type Operators numOutputRows number of output rows sum All aggregate operators aggTime time in aggregation build timing HashAggregateExec, ObjectHashAggregateExec, SortAggregateExec peakMemory peak memory size HashAggregateExec spillSize spill size size HashAggregateExec, ObjectHashAggregateExec avgHashProbe avg hash probes per key average HashAggregateExec numTasksFallBacked number of sort fallback tasks sum HashAggregateExec, ObjectHashAggregateExec Join Operators Metric Display Name Type Operators numOutputRows number of output rows sum All join operators buildDataSize data size of build side size ShuffledHashJoinExec buildTime time to build hash map timing ShuffledHashJoinExec spillSize spill size size SortMergeJoinExec Sort Operator Metric Display Name Type Operators sortTime sort time timing SortExec peakMemory peak memory size SortExec spillSize spill size size SortExec Shuffle Exchange Metric Display Name Type Operators dataSize data size size ShuffleExchangeExec numPartitions number of partitions sum ShuffleExchangeExec shuffleBytesWritten shuffle bytes written size Shuffle write shuffleRecordsWritten shuffle records written sum Shuffle write shuffleWriteTime shuffle write time nsTiming Shuffle write Shuffle Read (via AQEShuffleReadExec) Metric Display Name Type Operators numPartitions number of partitions sum AQEShuffleReadExec partitionDataSize partition data size size AQEShuffleReadExec numCoalescedPartitions number of coalesced partitions sum AQEShuffleReadExec numSkewedPartitions number of skewed partitions sum AQEShuffleReadExec numSkewedSplits number of skewed partition splits sum AQEShuffleReadExec numEmptyPartitions number of empty partitions sum AQEShuffleReadExec remoteBlocksFetched remote blocks read sum Shuffle read localBlocksFetched local blocks read sum Shuffle read remoteBytesRead remote bytes read size Shuffle read remoteBytesReadToDisk remote bytes read to disk size Shuffle read localBytesRead local bytes read size Shuffle read fetchWaitTime fetch wait time timing Shuffle read recordsRead records read sum Shuffle read remoteReqsDuration remote reqs duration timing Shuffle read remoteMergedReqsDuration remote merged reqs duration timing Shuffle read Broadcast Exchange Metric Display Name Type Operators dataSize data size size BroadcastExchangeExec numOutputRows number of output rows sum BroadcastExchangeExec collectTime time to collect timing BroadcastExchangeExec buildTime time to build timing BroadcastExchangeExec broadcastTime time to broadcast timing BroadcastExchangeExec Python UDF Operators Metric Display Name Type Operators pythonDataSent data sent to Python workers size All Python operators pythonDataReceived data returned from Python workers size All Python operators pythonBootTime time to start Python workers timing All Python operators pythonInitTime time to initialize Python workers timing All Python operators pythonTotalTime time to run Python workers timing All Python operators pythonProcessingTime time to execute Python code timing All Python operators pythonNumRowsReceived number of output rows sum All Python operators Window Operators Metric Display Name Type Operators spillSize spill size size WindowExec, ArrowWindowPythonExec Write Operators Metric Display Name Type Operators numFiles number of written files sum File writes numOutputBytes written output size File writes numOutputRows number of output rows sum File writes numParts number of dynamic part sum File writes taskCommitTime task commit time timing File writes jobCommitTime job commit time timing File writes MERGE INTO Operator Metric Display Name Type Operators numTargetRowsCopied target rows copied unmodified sum MergeRowsExec numTargetRowsInserted target rows inserted sum MergeRowsExec numTargetRowsUpdated target rows updated sum MergeRowsExec numTargetRowsDeleted target rows deleted sum MergeRowsExec numTargetRowsMatchedUpdated target rows updated by matched clause sum MergeRowsExec numTargetRowsMatchedDeleted target rows deleted by matched clause sum MergeRowsExec numTargetRowsNotMatchedBySourceUpdated target rows updated by not matched by source sum MergeRowsExec numTargetRowsNotMatchedBySourceDeleted target rows deleted by not matched by source sum MergeRowsExec Stateful Streaming Operators Metric Display Name Type Operators numOutputRows number of output rows sum Stateful operators numTotalStateRows number of total state rows sum Stateful operators numRowsDroppedByWatermark rows dropped by watermark sum Stateful operators stateMemory memory used by state size Stateful operators allUpdatesTimeMs time to update timing Stateful operators allRemovalsTimeMs time to remove timing Stateful operators commitTimeMs time to commit changes timing Stateful operators Other Operators Metric Display Name Type Operators numOutputRows number of output rows sum FilterExec, ProjectExec, ExpandExec, GenerateExec, CommandResultExec, WindowGroupLimitExec, UnionLoopExec, PythonWorkerLogsExec numAnchorOutputRows number of anchor output rows sum UnionLoopExec numIterations number of recursive iterations sum UnionLoopExec dataSize data size size SubqueryBroadcastExec collectTime time to collect timing SubqueryBroadcastExec Columnar Batch Operators Metric Display Name Type Operators numInputBatches number of input batches sum Columnar transitions numOutputBatches number of output batches sum Columnar transitions numInputRows number of input rows sum Columnar transitions numOutputRows number of output rows sum Columnar transitions WholeStageCodegen and Metric Scope Most operators (FilterExec, ProjectExec, HashAggregateExec, joins) are fused by WholeStageCodegen into a single JVM method. Their row count metrics (numOutputRows) are individually accurate, but they don\u0026rsquo;t have individual timing because they execute as one compiled function.\nOperators that have phases executing outside the codegen pipeline and have their own timing:\nSortExec (sort time) Aggregations (aggregation build time) ShuffledHashJoinExec (hash map build time) BroadcastExchangeExec (collect/build/broadcast time) ShuffleExchangeExec (shuffle write time) Python UDF operators (Python worker time) Stateful streaming operators (update/remove/commit time) In Part 2, we\u0026rsquo;ll cover how SQL metrics are implemented internally (the AccumulatorV2 lifecycle), and how AQE uses shuffle statistics at runtime to rewrite query plans. In Part 3, we\u0026rsquo;ll cover the DataSource V2 CustomMetric extension API, UI rendering, and the REST API.\n","permalink":"https://yaooqinn.github.io/posts/spark/understanding-sql-metrics/","summary":"Part 1 of a 3-part deep dive into Apache Spark\u0026rsquo;s SQL metrics system. Covers the 5 metric types, a complete reference of 100+ metrics across all operators, and how to read the numbers in the Spark UI.","title":"Deep Dive into Spark SQL Metrics (Part 1): Types, Full Reference, and What They Mean"},{"content":"This is Part 2 of a 3-part series on Spark SQL Metrics:\nPart 1: Metric types, complete reference, and what they mean Part 2 (this post): How metrics work internally, and how AQE uses them for runtime decisions Part 3: Extension APIs, UI rendering, and REST API Part 4: How Gluten extends the metrics system The AccumulatorV2 Lifecycle In Part 1, we looked at SQL metrics from the outside — what they measure and how to read the numbers. Now let\u0026rsquo;s trace how those numbers get from executor tasks to the Spark UI.\nFrom Task to Driver Every SQL metric is a SQLMetric, which extends AccumulatorV2[Long, Long]. When a physical operator defines a metric like numOutputRows, Spark creates an accumulator on the driver and registers it with the SparkContext.\nWhen a task runs on an executor, it works with a local copy of the accumulator. The operator code calls metric += value or metric.add(value) as it processes rows. These updates are purely local — no network traffic during execution.\nThe interesting part happens when the task finishes:\nTask (Executor) Driver ───────────── ────── metric.add(value) ↓ Task completes →────→ onTaskEnd() ↓ Store in LiveStageMetrics ↓ aggregateMetrics() ↓ MetricUtils.stringValue() ↓ Map[accId → \u0026#34;512.0 MiB (min, med, max)\u0026#34;] ↓ Persist to KVStore (SQLExecutionUIData) On task completion, the driver receives the accumulator updates through SparkListener events. The SQLAppStatusListener handles onTaskEnd() — it takes the metric values from the completed task and stores them in LiveStageMetrics, an in-memory structure that tracks per-task values for each stage.\nAggregation and Storage For completed executions, metrics go through aggregateMetrics(), which computes the total (min, med, max) distribution you see in the UI. These aggregated values are formatted into human-readable strings by MetricUtils.stringValue() and persisted to the KVStore as part of SQLExecutionUIData. Once stored, the original per-task values are discarded.\nFor live (still-running) executions, the aggregation happens on-the-fly. Each time you refresh the SQL tab, the listener computes the distribution from whatever task values are currently in memory. This is why metrics update in near-real-time while a query is running.\nDriver-Side Metrics Not all metrics come from tasks. Some originate on the driver itself:\nSubquery execution time — when a scalar subquery runs, the driver times it and posts the result Broadcast time — the time the driver spends broadcasting a table to executors These driver-side metrics use SQLMetrics.postDriverMetricUpdates(), which directly updates the accumulator on the driver without going through the task lifecycle. They bypass onTaskEnd() entirely.\nHow AQE Uses Statistics for Runtime Decisions This is where things get subtle — and where many people get confused. Adaptive Query Execution (AQE) makes smart decisions at runtime based on actual data sizes. But it doesn\u0026rsquo;t use SQL Metrics to do this. It uses a completely separate data source: MapOutputStatistics.\nThe Data Flow When AQE is enabled, Spark doesn\u0026rsquo;t execute the entire query plan at once. Instead, it executes stage by stage:\nShuffleExchangeExec submits a shuffle map stage via sparkContext.submitMapStage() The map stage runs — tasks write shuffle data to local disk After all map tasks complete, the MapOutputTracker knows exactly how many bytes each reducer partition will receive This information is packaged as MapOutputStatistics, which contains bytesByPartitionId: Array[Long] — the exact byte size of each shuffle partition ShuffleQueryStageExec exposes this via its mapStats property AdaptiveSparkPlanExec runs optimization rules after the stage materializes, using these real statistics instead of estimates The key insight: AQE waits for shuffle stages to finish, then uses the actual output sizes to decide what to do next.\nCoalesceShufflePartitions — Merging Small Partitions The most common AQE optimization. After a shuffle, you might have 200 partitions (the default spark.sql.shuffle.partitions) where most contain very little data.\nCoalesceShufflePartitions reads bytesByPartitionId and merges adjacent small partitions until each merged partition reaches approximately spark.sql.adaptive.advisoryPartitionSizeInBytes (default 64 MB).\nKey configuration:\nConfig Default Purpose spark.sql.adaptive.advisoryPartitionSizeInBytes 64 MB Target size for coalesced partitions spark.sql.adaptive.coalescePartitions.minPartitionNum (none) Minimum number of partitions to keep spark.sql.adaptive.coalescePartitions.minPartitionSize 1 MB Won\u0026rsquo;t create partitions smaller than this Example: If you have 200 partitions averaging 1 MB each, AQE might coalesce them into ~3 partitions of ~64 MB each. Instead of 200 tasks reading tiny amounts of data, you get 3 tasks doing real work.\nOptimizeSkewedJoin — Splitting Skewed Partitions Data skew is one of the most common performance problems in Spark. One partition has 10 GB while the rest have 100 MB each — the skewed partition becomes a bottleneck.\nOptimizeSkewedJoin reads bytesByPartitionId for both sides of a shuffle join. It calculates the median partition size, then flags a partition as \u0026ldquo;skewed\u0026rdquo; if:\nsize \u0026gt; max(skewThreshold, median × skewFactor) Key configuration:\nConfig Default Purpose spark.sql.adaptive.skewJoin.skewedPartitionThresholdInBytes 256 MB Absolute minimum to be considered skewed spark.sql.adaptive.skewJoin.skewedPartitionFactor 5.0 Must be this many times the median Both conditions must be met: the partition must be at least 256 MB and at least 5× the median size.\nOnce a skewed partition is identified, AQE splits it into smaller sub-partitions, each targeting advisoryPartitionSizeInBytes (64 MB). The non-skewed side of the join is duplicated to match — each sub-partition on the skewed side gets a full copy of the corresponding partition from the other side.\nOptimizeShuffleWithLocalRead — Eliminating Shuffle Network I/O When AQE determines that the shuffle data can be read locally (co-located on the same executor), it replaces the standard shuffle read with AQEShuffleReadExec configured for local reading. This eliminates network transfer entirely — the reducer reads shuffle files directly from local disk.\nThis optimization most commonly applies after broadcast hash joins (where all data is already local), but can also apply to other shuffles when the partitioning allows local reads.\nThe Distinction: SQL Metrics vs AQE Statistics This is the most important conceptual distinction in this post:\nSQL Metrics AQE Statistics What SQLMetric accumulators MapOutputStatistics Purpose Observability (what you see in the UI) Runtime plan optimization Data format Formatted strings (\u0026quot;512.0 MiB\u0026quot;) Raw Long[] arrays (byte counts) Code path AccumulatorV2 → SparkListener → KVStore MapOutputTracker → ShuffleQueryStageExec.mapStats When computed After each task completes After all map tasks in a stage complete Who consumes Spark UI, REST API, you AQE optimizer rules They often measure similar things — both care about data sizes — but through completely different code paths. SQL Metrics tell you what happened. AQE Statistics determine what happens next.\nThat said, AQE\u0026rsquo;s actions do show up in SQL Metrics. When AQE coalesces or splits partitions, the resulting AQEShuffleReadExec operator reports its own metrics that tell you exactly what AQE decided to do.\nMetrics That Tell You What AQE Did The AQEShuffleReadExec operator (covered in Part 1) is your window into AQE\u0026rsquo;s decisions. Here\u0026rsquo;s what each metric tells you:\nMetric What It Means numCoalescedPartitions \u0026gt; 0 AQE merged small partitions together numSkewedPartitions \u0026gt; 0 AQE detected skewed partitions numSkewedSplits How many sub-partitions were created from skewed ones numEmptyPartitions Empty partitions that were detected partitionDataSize Actual data size after AQE optimization Practical example: If you see numSkewedPartitions: 3 and numSkewedSplits: 12, it means AQE found 3 partitions that exceeded the skew threshold and split them into 12 sub-partitions. Those 3 bottleneck tasks became 12 parallel tasks, dramatically reducing wall-clock time.\nIf you see numCoalescedPartitions: 180 with the original numPartitions: 200, AQE merged 180 tiny partitions together — your 200 reducer tasks likely became around 20.\nThese metrics are the way to confirm whether AQE is actually helping your query. If numCoalescedPartitions and numSkewedPartitions are both zero, AQE is active but didn\u0026rsquo;t find anything to optimize.\nUsing SQL Plans to Understand AQE The SQL execution plan is another powerful tool for understanding what AQE did. When AQE is active, the plan shows AdaptiveSparkPlan at the top with isFinalPlan=true (for completed executions).\nYou can compare the initial plan (what the optimizer originally planned) with the final plan (what actually executed after AQE\u0026rsquo;s changes):\n# See the initial plan (before AQE) spark-history-cli -a \u0026lt;app-id\u0026gt; sql-plan \u0026lt;execution-id\u0026gt; --view initial # See the final plan (after AQE) spark-history-cli -a \u0026lt;app-id\u0026gt; sql-plan \u0026lt;execution-id\u0026gt; --view final By comparing these two plans, you can see exactly where AQE intervened:\nShuffleExchangeExec nodes replaced by AQEShuffleReadExec — shuffle optimizations applied Join strategy changes — e.g., sort-merge join converted to broadcast hash join because one side turned out to be small Different partition counts in the final plan — coalescing or splitting happened This comparison is invaluable when debugging performance issues: you can see whether AQE\u0026rsquo;s decisions helped or whether further tuning is needed.\nIn Part 1, we covered the five metric types and the complete reference. In Part 3, we\u0026rsquo;ll cover the DataSource V2 CustomMetric extension API, how the UI renders metrics, and how to query them programmatically via the REST API.\n","permalink":"https://yaooqinn.github.io/posts/spark/sql-metrics-part2-internals/","summary":"Part 2 of the SQL Metrics deep dive. How metrics flow from tasks to driver, and how Adaptive Query Execution uses shuffle statistics to rewrite plans at runtime.","title":"Deep Dive into Spark SQL Metrics (Part 2): Internals and How AQE Uses Them"},{"content":"This is Part 3 of a 3-part series on Spark SQL Metrics:\nPart 1: Metric types, complete reference, and what they mean Part 2: How metrics work internally, and how AQE uses them for runtime decisions Part 3 (this post): Extension APIs, UI rendering, and REST API Part 4: How Gluten extends the metrics system The DataSource V2 CustomMetric API In Part 1 and Part 2, we explored Spark\u0026rsquo;s built-in metrics and their internal machinery. But what if you\u0026rsquo;re building a custom connector and need to expose connector-specific numbers — bytes read from your proprietary format, cache hit rates, or throttling counts? Since Spark 3.2, the DataSource V2 API provides a clean extension point for exactly this.\nThe Interface Hierarchy At the core are two interfaces in org.apache.spark.sql.connector.metric:\nCustomMetric — defined once in your connector, describes what the metric is:\nname() — the metric name (must match between CustomMetric and CustomTaskMetric) description() — human-readable description shown in the UI aggregateTaskMetrics(long[] taskMetrics) — you decide how to combine per-task values into a single display string. This is where the connector has full control: you could compute a sum, average, percentile, or anything else. Must have a zero-argument constructor — Spark instantiates it via reflection on the driver when aggregating. CustomTaskMetric — reported by each PartitionReader on executors:\nname() — must match the corresponding CustomMetric.name() value() — returns a long representing the current metric value for this task Spark ships two convenient base classes so you don\u0026rsquo;t need to implement aggregateTaskMetrics from scratch:\nClass Aggregation Logic Output Format CustomSumMetric Sums all task values String.valueOf(sum) CustomAvgMetric Computes average of task values DecimalFormat(\u0026quot;#0.000\u0026quot;).format(avg) How to Implement Custom Metrics Step 1: Define your metric class.\nExtend one of the built-in base classes or implement CustomMetric directly:\npublic class MyBytesReadMetric extends CustomSumMetric { @Override public String name() { return \u0026#34;myBytesRead\u0026#34;; } @Override public String description() { return \u0026#34;bytes read from my source\u0026#34;; } } Step 2: Register the metric in your Scan.\nYour Scan implementation tells Spark which custom metrics your connector supports:\n@Override public CustomMetric[] supportedCustomMetrics() { return new CustomMetric[] { new MyBytesReadMetric() }; } Step 3: Report values from your PartitionReader.\nEach PartitionReader reports its current metric values whenever Spark calls currentMetricsValues(). This is called every 100 rows (controlled by CustomMetrics.NUM_ROWS_PER_UPDATE) and at task completion:\n@Override public CustomTaskMetric[] currentMetricsValues() { return new CustomTaskMetric[] { new CustomTaskMetric() { @Override public String name() { return \u0026#34;myBytesRead\u0026#34;; } @Override public long value() { return bytesReadSoFar; } } }; } That\u0026rsquo;s it — your custom metric will now appear in the Spark UI alongside the built-in ones.\nWrite-Side Custom Metrics Custom metrics aren\u0026rsquo;t limited to reads. Write connectors can also define custom metrics through BatchWrite.supportedCustomMetrics(), with values reported via DataWriter.currentMetricsValues(). This is useful for metrics like compression ratios, flush counts, or batching statistics on the write path.\nHow Spark Processes Custom Metrics Internally Behind the scenes, several components work together to make custom metrics flow through the same pipeline as built-in metrics:\nRegistration: DataSourceV2ScanExecBase calls scan.supportedCustomMetrics() during planning and creates SQLMetric wrappers via SQLMetrics.createV2CustomMetric(). Each wrapper gets a special type string.\nType encoding: The metric type is stored as \u0026quot;v2Custom_\u0026lt;fully.qualified.ClassName\u0026gt;\u0026quot; — for example, \u0026quot;v2Custom_com.mycompany.MyBytesReadMetric\u0026quot;. This encoding is constructed by CustomMetrics.buildV2CustomMetricTypeName().\nAggregation: When SQLAppStatusListener receives task metrics during aggregation, it parses the v2Custom_ prefix, extracts the class name, loads it via reflection, and calls aggregateTaskMetrics(long[]) on the instantiated object. This is why the zero-arg constructor is required.\nSpecial metric names: If your CustomTaskMetric uses the names \u0026quot;bytesWritten\u0026quot; or \u0026quot;recordsWritten\u0026quot;, Spark also propagates the values to its internal TaskOutputMetrics. This means they will appear in the Executors tab and stage-level I/O summaries, not just in the SQL tab.\nDriver-Side Custom Metrics Custom metrics aren\u0026rsquo;t limited to executor-side reporting. Your Scan can also report metrics from the driver:\nScan.reportDriverMetrics() returns a CustomTaskMetric[] array from the driver side DataSourceV2ScanExecBase.postDriverMetrics() posts them to the metrics system via SQLMetrics.postDriverMetricUpdates() This is useful for metrics like \u0026ldquo;number of files listed\u0026rdquo;, \u0026ldquo;partitions pruned\u0026rdquo;, or \u0026ldquo;metadata cache hits\u0026rdquo; — things that happen during planning on the driver rather than during data reading on executors.\nHow Metrics Are Rendered in the UI Once metrics are collected and aggregated on the driver, they need to be rendered. The Spark UI\u0026rsquo;s SQL tab has evolved significantly, and understanding the rendering pipeline helps you interpret what you see.\nThe Plan Visualization Pipeline The journey from stored metrics to visual rendering follows this path:\nSQLAppStatusStore.executionMetrics(id) → Map[accumulatorId → formatted String] ExecutionPage.planVisualization() → graph.makeDotFile(metrics) # compact DOT labels → graph.makeNodeDetailsJson(metrics) # full metrics JSON spark-sql-viz.js → renderPlanViz() # dagre-d3 graph → getNodeDetails() # parse JSON → updateDetailsPanel() # side panel on click → rerenderWithDetailedLabels() # optional inline mode The server side prepares two representations: a DOT file for the graph layout (with compact node labels) and a JSON payload with full metric details. The JavaScript frontend renders the DAG using dagre-d3 and provides interactive metric exploration.\nCompact vs Detailed Mode The SQL plan visualization supports two display modes:\nCompact mode (default since SPARK-55785): Node labels show only operator names. Metrics are available through clicking a node, which opens a side panel with the full metric table. This keeps the graph readable even for complex plans with dozens of operators.\nDetailed mode (toggle via checkbox): Metrics are rendered inline inside graph nodes in a 10px font. Useful when you want a printable snapshot of the full plan with all numbers, but can make the graph very wide for operators with many metrics.\nStage/Task toggle: When enabled, adds (stage X: task Y) annotations to max values, helping you identify which specific task produced the extreme value — invaluable for debugging skew.\nThe Side Panel When you click a node in compact mode, the side panel shows:\nMetric name + formatted value in a clean table layout WholeStageCodegen clusters: Clicking a cluster node shows all child operator metrics grouped together, so you can see the full picture of what happened inside a single codegen unit Search filter: For plans with many metrics, a text filter helps you quickly find the metric you care about Description tooltip: Hover over the operator name in the panel title to see a tooltip that helps disambiguate operators when the same operator type appears multiple times in a plan The REST API The Spark UI is great for visual exploration, but for automation — monitoring dashboards, performance regression tests, or post-hoc analysis scripts — you need programmatic access.\nEndpoint The primary endpoint for SQL execution metrics is:\nGET /api/v1/applications/{appId}/sql/{executionId} Query parameters:\nParameter Default Description details true Include node-level details with metrics planDescription true Include the physical plan text Response Structure A typical response looks like:\n{ \u0026#34;id\u0026#34;: 0, \u0026#34;status\u0026#34;: \u0026#34;COMPLETED\u0026#34;, \u0026#34;description\u0026#34;: \u0026#34;count at ...\u0026#34;, \u0026#34;planDescription\u0026#34;: \u0026#34;*(1) HashAggregate ...\u0026#34;, \u0026#34;submissionTime\u0026#34;: \u0026#34;2026-04-01T12:00:00Z\u0026#34;, \u0026#34;duration\u0026#34;: 5432, \u0026#34;runningJobIds\u0026#34;: [], \u0026#34;successJobIds\u0026#34;: [0, 1], \u0026#34;failedJobIds\u0026#34;: [], \u0026#34;nodes\u0026#34;: [ { \u0026#34;nodeId\u0026#34;: 0, \u0026#34;nodeName\u0026#34;: \u0026#34;HashAggregate\u0026#34;, \u0026#34;wholeStageCodegenId\u0026#34;: 1, \u0026#34;metrics\u0026#34;: [ {\u0026#34;name\u0026#34;: \u0026#34;number of output rows\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;5,000\u0026#34;}, {\u0026#34;name\u0026#34;: \u0026#34;peak memory\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;total (min, med, max)\\n512.0 MiB (128.0 MiB, 128.0 MiB, 128.0 MiB)\u0026#34;}, {\u0026#34;name\u0026#34;: \u0026#34;spill size\u0026#34;, \u0026#34;value\u0026#34;: \u0026#34;0.0 B\u0026#34;} ] } ], \u0026#34;edges\u0026#34;: [ {\u0026#34;fromId\u0026#34;: 1, \u0026#34;toId\u0026#34;: 0} ] } Key Things to Note Metric is just {name: String, value: String} — the REST API returns the formatted display string, not raw numeric values or metric types. If you need to do arithmetic on metric values, you\u0026rsquo;ll need to parse the formatted strings yourself.\nwholeStageCodegenId tells you which codegen cluster an operator belongs to. Operators with the same ID were fused into a single generated Java class.\nedges define the DAG structure as parent→child operator relationships. Combined with nodeId values, you can reconstruct the full plan tree programmatically.\nListing endpoint: To get all SQL executions for an application:\nGET /api/v1/applications/{appId}/sql?offset=0\u0026amp;length=100 Supports pagination via offset and length parameters.\nAccessing via spark-history-cli For interactive exploration, spark-history-cli wraps the REST API with convenient commands:\n# Structured JSON output spark-history-cli --json -a \u0026lt;app\u0026gt; sql # list all SQL executions spark-history-cli --json -a \u0026lt;app\u0026gt; sql \u0026lt;id\u0026gt; # single execution with metrics # Plan text spark-history-cli -a \u0026lt;app\u0026gt; sql-plan \u0026lt;id\u0026gt; # full plan spark-history-cli -a \u0026lt;app\u0026gt; sql-plan \u0026lt;id\u0026gt; --view final # post-AQE plan Practical Examples Using Spark Listener to Capture Metrics Programmatically If you want to react to metrics in real time — for example, logging slow queries or triggering alerts — you can register a QueryExecutionListener:\nspark.listenerManager.register(new QueryExecutionListener { override def onSuccess(funcName: String, qe: QueryExecution, durationNs: Long): Unit = { val metrics = qe.executedPlan.collectLeaves().flatMap(_.metrics) metrics.foreach { case (name, metric) =\u0026gt; println(s\u0026#34;$name: ${metric.value}\u0026#34;) } } override def onFailure(funcName: String, qe: QueryExecution, exception: Exception): Unit = {} }) This listener fires after every successful query execution and gives you access to the executed physical plan, where you can traverse operators and read their metric values directly as raw Long values — this is the only way to access metrics without the display formatting and rounding applied by the UI and REST API.\nAccessing Metrics from DataFrame Execution For ad-hoc debugging or REPL-based exploration, you can access metrics after executing a query through the status store:\nval df = spark.sql(\u0026#34;SELECT count(*) FROM my_table\u0026#34;) df.collect() // Access the last execution\u0026#39;s metrics val lastExec = spark.sharedState.statusStore.executionsList().last val metrics = spark.sharedState.statusStore.executionMetrics(lastExec.executionId) metrics.foreach { case (accId, value) =\u0026gt; println(s\u0026#34;$accId: $value\u0026#34;) } This approach is useful for integration tests where you want to assert that a specific optimization was applied (e.g., \u0026ldquo;number of files pruned\u0026rdquo; \u0026gt; 0) or for notebooks where you want to inspect performance without switching to the UI.\nNote: spark.sharedState.statusStore is internal API and only available on the driver side. In Spark Connect mode, clients don\u0026rsquo;t have access to the status store — use the REST API instead.\nSeries Conclusion This concludes our 3-part deep dive into Spark SQL Metrics:\nIn Part 1, we established the foundation: the five metric types (sum, size, timing, nanoTiming, average), the total (min, med, max) aggregation format, and a comprehensive reference of 100+ metrics across all operators.\nIn Part 2, we traced the internal lifecycle — how AccumulatorV2 values flow from executor tasks to the driver, how SQLAppStatusListener aggregates them, and how Adaptive Query Execution uses shuffle statistics (not SQL metrics) to make runtime decisions like partition coalescing, skew join optimization, and local shuffle reads.\nIn Part 3 (this post), we covered the extension points: how connector developers can define custom metrics via the DataSource V2 API, how the UI renders plans and metrics through the DOT/JSON/dagre-d3 pipeline, and how to query metrics programmatically via the REST API and Spark listeners.\nTogether, these three perspectives — what metrics measure, how they work internally, and how to extend and access them — give you the complete picture needed to effectively use SQL metrics for performance debugging, monitoring, and connector development.\nIn Part 1, we covered the five metric types and the complete reference. In Part 2, we traced the internal lifecycle and AQE\u0026rsquo;s use of shuffle statistics. This concludes the series.\n","permalink":"https://yaooqinn.github.io/posts/spark/sql-metrics-part3-extension-api/","summary":"Part 3 of the SQL Metrics deep dive. How to extend Spark with custom metrics via the DataSource V2 API, how the UI renders them, and how to query metrics programmatically.","title":"Deep Dive into Spark SQL Metrics (Part 3): Extension APIs, UI, and REST API"},{"content":"This is a bonus Part 4 of the Spark SQL Metrics series:\nPart 1: Metric types, complete reference, and what they mean Part 2: How metrics work internally, and how AQE uses them for runtime decisions Part 3: Extension APIs, UI rendering, and REST API Part 4 (this post): How Gluten extends the metrics system How Gluten\u0026rsquo;s Native Engine Produces Metrics Apache Gluten replaces the JVM execution engine with a native C++ engine — either Velox or ClickHouse. Because native operators execute independently (not fused by JVM codegen), each C++ operator is a separate function call with its own timing infrastructure. As a natural consequence, Gluten surfaces 60+ metrics per operator, including wall clock time, per-phase join metrics, native spill tracking, dynamic filter statistics, and I/O breakdowns by storage tier.\nThe 3-Layer Architecture Gluten\u0026rsquo;s metrics system bridges two worlds: Spark\u0026rsquo;s JVM-based SQLMetric framework and the native C++ execution engine. The architecture has three layers:\nSpark SQLMetric (JVM) ←── MetricsUpdater (bridge) ←── Velox/CH (C++) Map[String, SQLMetric] updateNativeMetrics() long[] arrays via JNI Layer 1: Spark SQLMetric (unchanged) Each *ExecTransformer — Gluten\u0026rsquo;s replacement for vanilla Spark\u0026rsquo;s *Exec operators — overrides lazy val metrics using the same pattern as vanilla Spark. But instead of hardcoding the metric set, it delegates to the backend:\nBackendsApiManager.getMetricsApiInstance .genFilterTransformerMetrics(sparkContext) This means the Velox backend and the ClickHouse backend can define completely different metrics for the same logical operator. A FilterExecTransformer running on Velox might expose wallNanos and peakMemoryBytes, while the same operator running on ClickHouse could expose different internal counters. The metric definitions are backend-specific, but they all end up as standard SQLMetric objects that Spark\u0026rsquo;s UI and REST API can display.\nLayer 2: MetricsUpdater (Gluten\u0026rsquo;s bridge abstraction) The MetricsUpdater trait is Gluten\u0026rsquo;s central bridging abstraction. It defines a single method:\ntrait MetricsUpdater extends Serializable { def updateNativeMetrics(opMetrics: IOperatorMetrics): Unit } Each operator has a corresponding MetricsUpdater implementation. These updaters are organized into a MetricsUpdaterTree that mirrors the plan DAG — one updater per operator, connected in the same parent-child structure as the physical plan.\nWhy a separate tree? Because the MetricsUpdaterTree is Serializable — it can be sent to executors without serializing the full SparkPlan (which contains non-serializable objects like SparkContext). On the executor, after native execution completes, the tree walks the native metrics and updates the SQLMetric accumulators.\nThree special sentinel instances handle edge cases:\nMetricsUpdater.None — operator has no metrics to update MetricsUpdater.Todo — metrics support not yet implemented for this operator MetricsUpdater.Terminate — the branch ends here (no children to recurse into) Here\u0026rsquo;s a concrete example — the FilterMetricsUpdater:\nclass FilterMetricsUpdater(val metrics: Map[String, SQLMetric]) extends MetricsUpdater { override def updateNativeMetrics(opMetrics: IOperatorMetrics): Unit = { val m = opMetrics.asInstanceOf[OperatorMetrics] metrics(\u0026#34;numOutputRows\u0026#34;) += m.outputRows metrics(\u0026#34;outputVectors\u0026#34;) += m.outputVectors metrics(\u0026#34;outputBytes\u0026#34;) += m.outputBytes metrics(\u0026#34;cpuCount\u0026#34;) += m.cpuCount metrics(\u0026#34;wallNanos\u0026#34;) += m.wallNanos metrics(\u0026#34;peakMemoryBytes\u0026#34;) += m.peakMemoryBytes metrics(\u0026#34;numMemoryAllocations\u0026#34;) += m.numMemoryAllocations } } Notice how each native metric field (e.g., m.wallNanos) maps directly to a SQLMetric key. The updater is the translation layer between native C++ naming and Spark\u0026rsquo;s metric namespace.\nLayer 3: Native metrics via JNI On the C++ side, the Velox engine collects metrics in arrays during execution — one entry per operator index. When a task completes, Gluten transfers these metrics across the JNI boundary as a Metrics object containing long[] arrays:\ninputRows[] — rows consumed by each operator outputRows[] — rows produced by each operator wallNanos[] — wall clock nanoseconds per operator cpuCount[] — CPU time per operator peakMemoryBytes[] — peak memory per operator ... — 20+ more arrays The MetricsUpdatingFunction walks the MetricsUpdaterTree, extracting per-operator values from the arrays by operator index. This is a bulk transfer — one JNI call per task, not per row — keeping overhead minimal.\nWhat Gluten Adds — 60+ Metrics Let\u0026rsquo;s look at the specific metrics Gluten introduces, organized by category.\nPer-Operator Execution Metrics In vanilla Spark, most operators report only numOutputRows. In Gluten, every operator gets these:\nMetric Display Name Type What It Measures wallNanos time of {operator} nsTiming Wall clock time per operator cpuCount cpu wall time count sum Number of getOutput() invocations (batch count) peakMemoryBytes peak memory bytes size Peak memory usage numMemoryAllocations number of memory allocations sum Memory allocation count outputRows number of output rows sum Output row count outputVectors number of output vectors sum Output vector (batch) count outputBytes number of output bytes size Output data volume in columnar format loadLazyVectorTime time to load lazy vectors timing Time loading lazy-evaluated vectors Note: wallNanos uses an operator-specific display name — \u0026ldquo;time of filter\u0026rdquo;, \u0026ldquo;time of sort\u0026rdquo;, \u0026ldquo;time of scan and filter\u0026rdquo;, \u0026ldquo;time of project\u0026rdquo;, etc.\nHaving wallNanos on every operator makes it straightforward to identify bottleneck operators in native execution.\nUnderstanding wallNanos and cpuCount These two metrics deserve special attention because they are the most important for performance analysis.\nBoth originate from Velox\u0026rsquo;s CpuWallTiming structure, which is collected via RAII timers (DeltaCpuWallTimer) wrapping each operator\u0026rsquo;s getOutput() call:\nstruct CpuWallTiming { uint64_t count; // Number of getOutput() invocations (batch count) uint64_t wallNanos; // Total wall-clock time (steady_clock, nanoseconds) uint64_t cpuNanos; // Total CPU time (CLOCK_THREAD_CPUTIME_ID, nanoseconds) }; wallNanos — measured with std::chrono::steady_clock. Captures total real elapsed time, including any time the operator spends blocked waiting for its child to produce data, I/O waits, or thread scheduling delays.\ncpuCount — despite the name, this is actually the invocation count (number of getOutput() calls = number of batches processed), not CPU time. The Gluten JNI bridge maps CpuWallTiming.count to the cpuCount metric.\nHow to interpret:\nScenario wallNanos cpuCount What It Means Large data, even work High High Many batches processed, expected Few batches, each slow High Low Possible skew or complex per-batch work Leaf operator (scan) High — Mostly I/O time (check ioWaitTime separately) Middle operator (filter) High — Includes wait for child — compare with child\u0026rsquo;s wallNanos Important caveat — wallNanos includes child waiting:\nBecause wallNanos wraps the entire getOutput() call, a parent operator\u0026rsquo;s wallNanos includes time spent blocked waiting for its child to produce data. This means:\nFor a leaf operator (scan): wallNanos ≈ I/O + compute time For a middle operator (filter above a scan): wallNanos = own compute + child\u0026rsquo;s scan time You cannot simply sum wallNanos across all operators — that would double-count To isolate an operator\u0026rsquo;s own contribution, compare its wallNanos with its child\u0026rsquo;s wallNanos. The difference is the operator\u0026rsquo;s own processing time. Velox also tracks some I/O-specific metrics separately (ioWaitTime, dataSourceReadTime) to help separate pure I/O from compute.\nScan-Specific Metrics Vanilla Spark\u0026rsquo;s scan operators have scanTime and numFiles. Gluten goes much deeper:\nMetric Display Name What It Measures skippedSplits / processedSplits number of skipped/processed splits File split pruning effectiveness skippedStrides / processedStrides number of skipped/processed row groups Row group/stripe pruning within files ioWaitTime io wait time Time waiting for I/O operations storageReadBytes storage read bytes Bytes read from remote storage localReadBytes Bytes read from local SSD cache ramReadBytes Bytes read from in-memory cache preloadSplits Pre-loaded splits (prefetching) dataSourceAddSplitTime Time managing split assignments dataSourceReadTime Time reading data from the source The storageReadBytes / localReadBytes / ramReadBytes breakdown is particularly valuable for cloud environments. If you see most reads coming from storageReadBytes, your cache isn\u0026rsquo;t warm. If ioWaitTime dominates wallNanos, the bottleneck is network I/O, not CPU.\nSpill Metrics Vanilla Spark tracks spill at the stage level. Gluten tracks it per operator, per phase:\nMetric Display Name What It Measures spilledBytes bytes written for spilling Volume of data spilled to disk spilledRows total rows written for spilling Number of rows spilled spilledPartitions total spilled partitions Number of partitions involved in spill spilledFiles total spilled files Number of spill files created For join operators, spill is tracked separately for the build and probe phases (see next section), so you can pinpoint exactly which phase is under memory pressure.\nDynamic Filter Metrics Dynamic filters (also called runtime filters) are generated by join operators to prune scan results at runtime. Vanilla Spark has no metrics for this. Gluten tracks the full lifecycle:\nMetric Display Name What It Measures numDynamicFiltersProduced number of dynamic filters produced Runtime filters generated by join build sides numDynamicFiltersAccepted number of dynamic filters accepted Runtime filters applied to scan operators numReplacedWithDynamicFilterRows number of replaced with dynamic filter rows Rows eliminated before reaching the join If numDynamicFiltersProduced \u0026gt; 0 but numDynamicFiltersAccepted = 0, the filters were generated but not applied — a sign that the scan and join aren\u0026rsquo;t connected in the way the optimizer expected. If numReplacedWithDynamicFilterRows is a large number, runtime filters are saving significant work.\nJoin Phase Separation — 20+ Metrics per Join This is arguably Gluten\u0026rsquo;s most powerful metric enhancement. Vanilla Spark\u0026rsquo;s join operators report a single buildTime and numOutputRows. Gluten splits every join into its constituent phases with separate metrics for each:\nBuild phase:\nMetric Display Name What It Measures hashBuildInputRows number of hash build input rows Rows consumed by the build side hashBuildOutputRows number of hash build output rows Rows in the hash table hashBuildWallNanos time of hash build Wall clock time for building hashBuildPeakMemoryBytes hash build peak memory bytes Peak memory during build hashBuildSpilledBytes hash build spilled bytes Data spilled during build hashBuildSpilledRows hash build spilled rows Rows spilled during build hashBuildSpilledPartitions hash build spilled partitions Partitions spilled during build hashBuildSpilledFiles hash build spilled files Spill files created during build Probe phase:\n| Metric | What It Measures |\nMetric Display Name What It Measures hashProbeInputRows number of hash probe input rows Rows consumed by the probe side hashProbeOutputRows number of hash probe output rows Rows output after probing hashProbeWallNanos time of hash probe Wall clock time for probing hashProbePeakMemoryBytes hash probe peak memory bytes Peak memory during probe hashProbeSpilledBytes hash probe spilled bytes Data spilled during probe hashProbeSpilledRows hash probe spilled rows Rows spilled during probe hashProbeSpilledPartitions hash probe spilled partitions Partitions spilled during probe hashProbeSpilledFiles hash probe spilled files Spill files created during probe Pre/post projection:\nMetric Display Name What It Measures streamPreProjectionWallNanos time of stream preProjection Expression evaluation time on the stream (probe) side before join streamPreProjectionCpuCount stream preProject cpu wall time count Batch count for stream pre-projection buildPreProjectionWallNanos time to build preProjection Expression evaluation time on the build side before join buildPreProjectionCpuCount preProject cpu wall time count Batch count for build pre-projection postProjectionWallNanos time of postProjection Expression evaluation time after join postProjectionCpuCount postProject cpu wall time count Batch count for post-projection In vanilla Spark, a slow join gives you almost nothing to work with — you know it\u0026rsquo;s slow, but not why. With Gluten, you can immediately see: is the build phase slow (maybe the build side is too large)? Is the probe phase slow (maybe hash collisions are causing excessive probing)? Is the build phase spilling (memory pressure)? This level of detail changes how you diagnose join performance.\nWrite Metrics Metric Display Name What It Measures physicalWrittenBytes number of written bytes Actual bytes written to storage writeIOTime / writeIONanos time of write IO I/O time during writes numWrittenFiles number of written files Number of files produced Reading Gluten Metrics in the Spark UI Gluten metrics appear in the same Spark SQL tab because they use the same SQLMetric framework. The operator names change (e.g., HashAggregateExecTransformer instead of HashAggregateExec) but metrics appear in the same side panel when you click on an operator node.\nWhat to Look For Here are the key patterns to watch for when reading Gluten metrics:\nIdentify the bottleneck operator:\nLook at wallNanos on each operator. In a healthy query, scan and join operators dominate. If a FilterExecTransformer or ProjectExecTransformer has high wallNanos, the filter or projection expression itself is expensive — consider simplifying it.\nDiagnose slow joins:\nCompare hashBuildWallNanos vs hashProbeWallNanos. If the build side dominates, the build input is too large — consider changing the join order or adding a filter to reduce the build side. If the probe side dominates, look at hashProbeInputRows — too many probe rows or hash collisions could be the cause.\nCheck native predicate pushdown:\nIf skippedSplits \u0026gt; 0, native file-level pruning is working. If skippedStrides \u0026gt; 0, row group or stripe-level pruning within files is working. If both are zero, your predicate isn\u0026rsquo;t being pushed down into the native scan — check if the column type supports pushdown.\nVerify runtime filter effectiveness:\nIf numDynamicFiltersAccepted \u0026gt; 0, runtime filters from join build sides are being applied to scans. Check numReplacedWithDynamicFilterRows to see how many rows were eliminated — a large number means significant I/O savings.\nDetect memory pressure in native engine:\nIf spilledBytes \u0026gt; 0 on any operator, the native engine is spilling to disk. For joins, check whether the build phase or probe phase is spilling. For aggregations, spill means the grouping cardinality is high. Consider increasing native memory allocation or reducing data volume.\nI/O tier analysis:\nCompare storageReadBytes, localReadBytes, and ramReadBytes on scan operators. In a well-cached environment, you want most reads from ramReadBytes or localReadBytes. High storageReadBytes means you\u0026rsquo;re reading from remote storage (S3, HDFS) — check if your caching layer is configured correctly.\nAccessing via spark-history-cli Gluten metrics are also available through the REST API and spark-history-cli, since they\u0026rsquo;re stored as standard SQLMetric values:\nspark-history-cli --json -a \u0026lt;app\u0026gt; sql \u0026lt;id\u0026gt; # includes Gluten metrics The JSON output will contain all the Gluten-specific metrics alongside vanilla Spark metrics, using the same {name, value} format described in Part 3.\nArchitectural Implications Gluten\u0026rsquo;s metrics system offers several insights about extending Spark\u0026rsquo;s observability:\nEngine replacement provides comprehensive metrics naturally. When the engine controls every operator\u0026rsquo;s execution, it can measure every boundary. Each C++ operator is a separate function call with its own start and end timestamps — per-operator timing on every operator is achievable without any workarounds.\nThe MetricsUpdater pattern is reusable. Any native backend can adopt this pattern: define a tree of lightweight, serializable updater objects that mirror the plan, transfer bulk metric arrays via JNI, and walk the tree to update SQLMetric accumulators.\nJNI array-based transfer minimizes overhead. Instead of calling back into the JVM for every metric update, Gluten batches all metrics into long[] arrays — one bulk JNI transfer per task. This keeps the metrics overhead negligible even with 60+ metrics per operator.\nBackend-agnostic design through MetricsApi. The MetricsApi abstraction means the Velox backend and ClickHouse backend can define completely different metrics for the same operator type. Adding a new backend (say, DataFusion) would only require implementing the MetricsApi interface — no changes to the core bridging code.\nIn Part 1, we covered the five metric types and the complete reference. In Part 2, we traced the internal lifecycle and AQE\u0026rsquo;s use of shuffle statistics. In Part 3, we explored extension APIs, UI rendering, and the REST API. This bonus Part 4 examined how Apache Gluten extends the metrics system by bridging native engine metrics back to Spark\u0026rsquo;s framework.\n","permalink":"https://yaooqinn.github.io/posts/spark/sql-metrics-part4-gluten/","summary":"Part 4 of the SQL Metrics deep dive. How Apache Gluten bridges native Velox/ClickHouse metrics back to Spark\u0026rsquo;s SQL Metrics framework, adding 60+ metrics that vanilla Spark doesn\u0026rsquo;t have.","title":"Deep Dive into Spark SQL Metrics (Part 4): How Gluten Extends the Metrics System"},{"content":"From DAGs to Declarations Every data engineer knows the drill: define Task A, Task B, Task C, draw the arrows, manage the DAG. Airflow does it. Dagster does it. Every orchestrator does it.\nBut here\u0026rsquo;s the thing: do you actually care about the DAG, or do you care about your data?\nApache Spark 4.1 offers a new answer: Spark Declarative Pipelines (SDP). You declare what tables you want and how their contents are derived. The framework handles dependency resolution, execution order, parallelization, error handling, and incremental updates.\nThis isn\u0026rsquo;t a small feature. It\u0026rsquo;s a fundamental rethinking of how data pipelines should be built.\nThree-Minute Quick Start Install:\npip install pyspark[pipelines] Write a pipeline:\nfrom pyspark import pipelines as dp @dp.materialized_view def daily_sales(): return spark.table(\u0026#34;orders\u0026#34;).groupBy(\u0026#34;date\u0026#34;).agg({\u0026#34;amount\u0026#34;: \u0026#34;sum\u0026#34;}) Run it:\nspark-pipelines run No saveAsTable(). No start(). No awaitTermination(). You described \u0026ldquo;I want a daily sales summary table\u0026rdquo; and SDP makes it exist.\nCore Concepts Flows: The Atomic Unit A flow describes a complete data movement: where to read, how to transform, where to write.\nStreaming Flow → outputs to a Streaming Table (incremental) Batch Flow → outputs to a Materialized View or Temporary View Datasets: What You Actually Care About Streaming Table — continuously updated from sources like Kafka:\n@dp.table def raw_events(): return ( spark.readStream.format(\u0026#34;kafka\u0026#34;) .option(\u0026#34;kafka.bootstrap.servers\u0026#34;, \u0026#34;localhost:9092\u0026#34;) .option(\u0026#34;subscribe\u0026#34;, \u0026#34;events\u0026#34;) .load() ) Materialized View — precomputed batch table, fully refreshed:\n@dp.materialized_view def hourly_metrics(): return ( spark.table(\u0026#34;raw_events\u0026#34;) .groupBy(window(\u0026#34;timestamp\u0026#34;, \u0026#34;1 hour\u0026#34;)) .agg(count(\u0026#34;*\u0026#34;).alias(\u0026#34;event_count\u0026#34;)) ) Temporary View — intermediate results scoped to one pipeline run:\n@dp.temporary_view def cleaned_events(): return spark.table(\u0026#34;raw_events\u0026#34;).filter(\u0026#34;event_type IS NOT NULL\u0026#34;) Automatic Dependency Inference You don\u0026rsquo;t declare dependencies. SDP analyzes your queries, discovers spark.table(\u0026quot;raw_events\u0026quot;) calls, and builds the dependency graph automatically:\nraw_events (Streaming Table) ↓ cleaned_events (Temporary View) ↓ hourly_metrics (Materialized View) SQL-Native Support The same pipeline in pure SQL:\nCREATE STREAMING TABLE raw_events AS SELECT * FROM STREAM kafka_source; CREATE TEMPORARY VIEW cleaned_events AS SELECT * FROM raw_events WHERE event_type IS NOT NULL; CREATE MATERIALIZED VIEW hourly_metrics AS SELECT window(timestamp, \u0026#39;1 hour\u0026#39;), count(*) AS event_count FROM cleaned_events GROUP BY 1; Zero learning curve for SQL-first teams.\nMixed Batch and Streaming Traditionally, batch and streaming are separate pipelines. SDP lets you mix them in one graph:\n@dp.table def orders(): return spark.readStream.format(\u0026#34;kafka\u0026#34;)... @dp.materialized_view def daily_summary(): return spark.table(\u0026#34;orders\u0026#34;).groupBy(\u0026#34;date\u0026#34;).count() SDP manages triggers, scheduling, and checkpoints automatically.\nEngineering: Project Structure and CLI spark-pipelines init --name my_pipeline # scaffold spark-pipelines dry-run # validate without I/O spark-pipelines run # execute Configure via spark-pipeline.yml:\nname: my_pipeline libraries: - glob: include: transformations/** catalog: my_catalog database: my_db The dry-run command catches syntax errors, analysis errors, and cyclic dependencies without touching any data — perfect for CI/CD.\nRelationship with Orchestrators SDP doesn\u0026rsquo;t replace Airflow or Dagster. It handles Spark-level data transformations and dependency management. In production:\nAirflow/Dagster (top-level orchestration) ├── Trigger SDP pipeline (data transformations) ├── Call external APIs ├── Send notifications └── Non-Spark tasks My Take as a Spark PMC Member SDP solves several long-standing pain points:\nLower barrier to entry. No need to understand checkpoint, trigger, outputMode to build reliable streaming pipelines.\nLess boilerplate. No more writeStream.format().option().start().awaitTermination() ceremony.\nUnified batch and streaming. Same declarative API, same dependency graph, no more two worlds.\nAI-friendly. Declarative flows are essentially functions — testable, callable, and easy for AI assistants to understand and generate.\nSDP\u0026rsquo;s design originates from Databricks\u0026rsquo; production-proven Delta Live Tables pattern, now brought to open-source Spark. The entire community benefits from best practices validated at massive scale.\nTry It pip install pyspark[pipelines] spark-pipelines init --name hello_sdp cd hello_sdp spark-pipelines run Full guide: Spark Declarative Pipelines Programming Guide\nSpark Declarative Pipelines was introduced in Apache Spark 4.1. Design doc: SPARK-51727.\n","permalink":"https://yaooqinn.github.io/posts/spark/spark-declarative-pipelines/","summary":"Apache Spark 4.1 introduces Spark Declarative Pipelines (SDP) — a declarative framework that lets you define what your data should look like, not how to compute it. As a Spark PMC Member, here\u0026rsquo;s my take on what this means for data engineering.","title":"Spark Declarative Pipelines: A Paradigm Shift for Data Engineering"},{"content":"Every Spark engineer has been there. A job that ran fine yesterday is now 3x slower. Someone asks \u0026ldquo;why is this query slow?\u0026rdquo; and you spend the next hour clicking through the Spark History Server UI, eyeballing numbers across tabs, mentally diffing configurations.\nWhat if your AI assistant could do that for you?\nWhat is spark-advisor? spark-advisor is an agent skill that turns your AI coding assistant into a Spark performance engineer. It works with GitHub Copilot, Claude Code, Cursor, and 30+ other agents.\nWhen you say \u0026ldquo;why is my Spark app slow?\u0026rdquo; or \u0026ldquo;compare these two TPC-DS runs\u0026rdquo;, the agent:\nConnects to your Spark History Server via spark-history-cli Collects structured JSON data (summary, stages, executors, SQL plans) Applies diagnostic heuristics to find bottlenecks Produces a prioritized report with actionable recommendations No manual clicking. No context-switching. Just ask.\nQuick Start Install the CLI and skill:\npip install spark-history-cli npx skills add yaooqinn/spark-history-cli Then tell your agent:\n\u0026ldquo;Diagnose the latest Spark app on my History Server\u0026rdquo;\nThat\u0026rsquo;s it. The agent will list apps, pick the latest, collect metrics, analyze them, and tell you what\u0026rsquo;s wrong.\nWhat It Diagnoses The skill encodes the diagnostic instincts of an experienced Spark engineer into structured rules. Here\u0026rsquo;s what it checks:\nTask Skew The single most common Spark performance killer. spark-advisor fetches task metric quantiles and compares p50 vs p95:\np95/p50 \u0026gt; 3x → moderate skew p95/p50 \u0026gt; 10x → severe skew It then recommends: AQE skew join, partition count tuning, or key salting — depending on the root cause.\nGC Pressure When GC time exceeds 10% of total executor runtime, spark-advisor flags it. At 20%+, it\u0026rsquo;s marked severe. Recommendations range from increasing executor memory to reducing per-executor concurrency.\nShuffle Overhead Shuffle-heavy stages are detected when shuffle bytes exceed 2x the input size. The skill checks for redundant exchanges in the plan, wrong join strategies (SortMergeJoin where BroadcastHashJoin would work), and insufficient partition counts.\nMemory Spill Any non-zero memoryBytesSpilled or diskBytesSpilled triggers an alert. Spill means the executor ran out of memory for aggregation or sort buffers — a problem that\u0026rsquo;s invisible in the UI unless you know where to look.\nStraggler Tasks Tasks that take 5x+ longer than the median in their stage. spark-advisor checks if the stragglers are on specific executors (hardware issue) or processing more data (skew variant).\nGluten/Velox Awareness For Gluten-accelerated Spark, the skill detects:\nFallback operators: Non-Transformer nodes in the final plan (e.g., SortMergeJoin instead of ShuffledHashJoinExecTransformer) Columnar-to-row transitions: VeloxColumnarToRow boundaries that indicate fallback Native metric patterns: Different GC and memory profiles in Gluten vs vanilla stages TPC-DS Benchmark Comparison One of spark-advisor\u0026rsquo;s most powerful features is structured benchmark comparison. Say:\n\u0026ldquo;Compare these two TPC-DS runs: app-20260315120000-0001 and app-20260320120000-0001\u0026rdquo;\nThe agent will:\nMatch queries across runs (q1–q99, handling split queries like q14a/b, q23a/b) Calculate per-query speedup and regression Produce a comparison table: | Query | Baseline | Candidate | Delta | Speedup | Status | |-------|----------|-----------|--------|---------|--------------| | q67 | 72s | 85s | +13s | 0.85x | ⚠ REGRESSED | | q1 | 61s | 45s | -16s | 1.36x | ✓ IMPROVED | | q50 | 34s | 33s | -1s | 1.03x | ≈ NEUTRAL | Drill into the top-3 regressions — comparing final plans, stage metrics, and config diffs Report overall speedup (geometric mean across all queries) This replaces hours of manual spreadsheet work after a benchmark run.\nThe Diagnostic Report spark-advisor produces a structured Markdown report:\n# Spark Performance Report ## Executive Summary The application spent 65% of total time in 3 shuffle-heavy stages. GC pressure is moderate (12% of executor time). Task skew detected in stage 14 (p95/p50 = 8.2x). ## Findings ### Finding 1: Severe Task Skew in Stage 14 - **Severity**: High - **Evidence**: p95 duration 45s vs p50 5.5s (8.2x ratio) - **Recommendation**: Enable AQE skew join optimization ### Finding 2: GC Pressure on Executors 3, 7 - **Severity**: Medium - **Evidence**: 18% GC time (threshold: 10%) - **Recommendation**: Increase spark.executor.memory from 4g to 8g ## Recommendations 1. Enable spark.sql.adaptive.skewJoin.enabled=true 2. Increase executor memory to 8g 3. Review join strategy for stage 14 — consider broadcast join How It Works Under the Hood spark-advisor is a pure SKILL.md — no code, just structured instructions that teach the agent how to be a Spark performance engineer. It uses:\nspark-history-cli for all data collection (--json mode for structured output) Diagnostic rules in references/diagnostics.md with specific thresholds and heuristics Comparison methodology in references/comparison.md for TPC-DS benchmarks Sample scripts in sample_codes/ for common patterns The agent reads these references, applies the rules to your data, and reasons about the results. No model fine-tuning, no training data — just well-structured domain knowledge.\nInstallation # Install the CLI pip install spark-history-cli # Install skills for your agent npx skills add yaooqinn/spark-history-cli This installs two skills:\nspark-history-cli — Query the Spark History Server spark-advisor — Diagnose, compare, and optimize What\u0026rsquo;s Next Auto-remediation: Suggest config patches that can be applied directly Historical trending: Track performance across releases Visualization: Generate SVG flamecharts and DAG annotations spark-advisor is open source at github.com/yaooqinn/spark-history-cli. Star it if you find it useful.\n","permalink":"https://yaooqinn.github.io/posts/spark/spark-advisor/","summary":"spark-advisor is an agent skill that turns your AI coding assistant into a Spark performance engineer — diagnosing slow jobs, detecting skew, comparing benchmark runs, and producing actionable tuning recommendations.","title":"Introducing spark-advisor: An AI-Powered Spark Performance Engineer"},{"content":"The Spark History Server has a decent web UI and a comprehensive REST API. But if you\u0026rsquo;re already in the terminal — SSH\u0026rsquo;d into a gateway node, debugging a pipeline in CI, or scripting a post-mortem — switching to a browser feels like a context switch you shouldn\u0026rsquo;t need to make.\nspark-history-cli puts the entire Spark History Server at your fingertips. It\u0026rsquo;s a Python CLI that wraps all 20 REST API endpoints into an interactive REPL and one-shot commands. List applications, drill into jobs and stages, inspect SQL executions, check executor stats, download event logs — all from your terminal.\nInstall pip install spark-history-cli That\u0026rsquo;s it. Requires Python 3.10+ and a running Spark History Server.\nTwo Modes of Operation Interactive REPL Just run spark-history-cli to enter the REPL:\n$ spark-history-cli --server http://my-shs:18080 spark-history\u0026gt; apps --status completed --limit 5 ID Name Status Start Time Duration app-20260318091500-0003 ETL Pipeline COMPLETED 2026-03-18 09:15:00 4m 32s app-20260318080000-0002 Daily Report COMPLETED 2026-03-18 08:00:00 12m 15s ... spark-history\u0026gt; use app-20260318091500-0003 Current app: app-20260318091500-0003 (ETL Pipeline) spark-history\u0026gt; jobs Job ID Status Stages Duration Description 0 SUCCEEDED 3/3 1m 02s save at ETLPipeline.scala:45 1 SUCCEEDED 2/2 2m 18s save at ETLPipeline.scala:78 2 SUCCEEDED 1/1 1m 12s save at ETLPipeline.scala:112 spark-history\u0026gt; stages spark-history\u0026gt; sql spark-history\u0026gt; executors spark-history\u0026gt; env The use command sets a \u0026ldquo;current app\u0026rdquo; context, so you don\u0026rsquo;t have to repeat the app ID on every command. It works exactly like USE database in SQL.\nOne-Shot Commands For scripting, CI pipelines, or quick lookups:\n# List completed apps spark-history-cli apps --status completed --limit 10 # Check jobs for a specific app spark-history-cli --app-id app-20260318091500-0003 jobs # Download event logs for offline analysis spark-history-cli --app-id app-20260318091500-0003 logs ./events.zip # JSON output for piping into jq or other tools spark-history-cli --json --app-id app-20260318091500-0003 stages The --json flag outputs raw JSON — perfect for piping into jq, feeding into monitoring scripts, or integrating with other tools.\nWhat You Can Do The CLI covers all 20 endpoints of the Spark History Server REST API:\nCommand What It Does apps List all applications with status, time, duration app \u0026lt;id\u0026gt; Show application details and set as current jobs List jobs with status, stages, and duration job \u0026lt;id\u0026gt; Show detailed job info stages List all stages stage \u0026lt;id\u0026gt; Show stage details with task summary executors List active executors executors --all Include dead executors sql List SQL executions sql \u0026lt;id\u0026gt; Show SQL execution details with plan graph rdds List cached RDDs env Show Spark configuration and environment logs [path] Download event logs as ZIP version Show History Server Spark version Why Not Just Use the Web UI? The web UI is great when you\u0026rsquo;re sitting at a browser. But there are real scenarios where a CLI is better:\nSSH debugging. You\u0026rsquo;re on a jump host or gateway node troubleshooting a production cluster. No browser, no port forwarding — just a terminal. spark-history-cli --server http://shs:18080 apps gets you started immediately.\nScripting and automation. Want to check if yesterday\u0026rsquo;s ETL jobs all succeeded? Write a cron job that runs spark-history-cli --json apps --status failed and alerts on non-empty output. The --json flag makes this trivial.\nPost-mortem workflows. Download event logs with logs, cross-reference job durations with jobs, check executor memory with executors --all — all in one terminal session without clicking through multiple browser tabs.\nCI/CD integration. After submitting a Spark application in a pipeline, query the History Server to verify the job completed successfully, check stage metrics, or archive event logs as build artifacts.\nWhy CLI, Not Just the Web UI? The Agentic Perspective The most important reason for spark-history-cli isn\u0026rsquo;t human convenience — it\u0026rsquo;s that AI agents can\u0026rsquo;t use web UIs.\nWe\u0026rsquo;re entering an era where LLM-powered agents — GitHub Copilot, coding assistants, on-call bots, automated root-cause analyzers — are becoming first-class participants in engineering workflows. These agents interact with the world through text interfaces: shell commands, APIs, and structured output. A web UI is a dead end for them. No matter how polished the Spark History Server\u0026rsquo;s web pages are, an agent can\u0026rsquo;t click links, scroll tables, or read DAG visualizations.\nA CLI changes everything:\nAgents can invoke it as a tool. When an agent needs to answer \u0026ldquo;why did last night\u0026rsquo;s ETL fail?\u0026rdquo;, it can run spark-history-cli --json apps --status failed, parse the JSON, pick the relevant app, run spark-history-cli --json --app-id \u0026lt;id\u0026gt; jobs to find the failed job, then stages to pinpoint the failing stage — all autonomously, in a chain-of-thought loop. The web UI offers no equivalent entry point for programmatic reasoning.\nStructured output enables reasoning. The --json flag isn\u0026rsquo;t just for jq — it\u0026rsquo;s what makes the tool legible to an LLM. An agent can ingest a JSON array of jobs, compare durations, spot anomalies, and synthesize a human-readable diagnosis. Try doing that with an HTML table rendered in a browser.\nThe REPL maps to how agents think. An agent exploring a Spark application follows the same drill-down pattern a human does: list apps → pick one → check jobs → drill into the slow stage → look at task metrics. The REPL\u0026rsquo;s use command and hierarchical navigation mirror this reasoning pattern naturally. Each command is a discrete, composable step an agent can plan and execute.\nIt completes the feedback loop. Consider a CI pipeline that submits a Spark application. Today, verifying the result means either parsing raw REST API responses with custom scripts or having a human check the web UI. With spark-history-cli, an agent (or a simple shell script) can query the History Server, verify success, extract metrics, and report — closing the automation loop entirely.\nThis is the real argument: the Spark History Server stores rich diagnostic data, but it\u0026rsquo;s locked behind a human-only interface. spark-history-cli turns that data into something both humans and agents can consume. In a world where your on-call assistant is an LLM, that distinction matters.\nGitHub Copilot CLI Skill spark-history-cli ships as a GitHub Copilot CLI skill — the agentic integration in practice. Install it with:\nspark-history-cli install-skill This copies the bundled skill definition to ~/.copilot/skills/spark-history-cli. After reloading skills (/skills reload), you can use natural language prompts like:\nUse /spark-history-cli to inspect the latest completed SHS application. Copilot CLI will invoke the tool, interpret the output, and answer your questions about Spark application history in conversational English. You describe the intent; the agent figures out which commands to run, chains them together, and synthesizes the answer. No command syntax to remember, no manual JSON parsing — just a question and an answer grounded in real History Server data.\nConfiguration The server URL defaults to http://localhost:18080. Override it with:\n# CLI flag spark-history-cli --server http://my-shs:18080 # Environment variable export SPARK_HISTORY_SERVER=http://my-shs:18080 spark-history-cli # REPL command (change on the fly) spark-history\u0026gt; server http://another-shs:18080 Get Started pip install spark-history-cli spark-history-cli The source is on GitHub: yaooqinn/spark-history-cli. It\u0026rsquo;s Apache 2.0 licensed. Issues, PRs, and feedback are welcome.\nspark-history-cli v1.0.1 is available on PyPI. Source code at github.com/yaooqinn/spark-history-cli.\n","permalink":"https://yaooqinn.github.io/posts/spark/spark-history-cli/","summary":"spark-history-cli brings the Spark History Server to your terminal — an interactive REPL and one-shot CLI that covers all 20 REST API endpoints. List apps, inspect jobs, drill into stages, check SQL executions, and download event logs without ever opening a browser. It also ships as a GitHub Copilot CLI skill.","title":"spark-history-cli: Making the Spark History Server Agent-Friendly"},{"content":"You click into a SQL execution in the Spark Web UI. You want to know: which jobs did this query launch, how far along are they, and did anything fail?\nHere\u0026rsquo;s what the page used to show you:\nRunning Jobs: 0, 1, 2 Succeeded Jobs: 3, 4 That\u0026rsquo;s it. Bare IDs. No status, no duration, no stage counts, no progress bars. To understand what was actually happening, you had to click each job ID, inspect the job detail page, navigate back, click the next one, and mentally assemble the picture. For a complex query that spawns a dozen jobs, this was a real pain.\nSPARK-55971 fixes this. The SQL execution detail page now has a full Associated Jobs table that shows everything at a glance.\nWhat It Looks Like Here\u0026rsquo;s the new jobs table on a succeeded execution:\nAnd here\u0026rsquo;s a failed execution with killed tasks — notice the concise progress bar labels:\nThe full page showing where the jobs table sits — right below Plan Details:\nFor comparison, here\u0026rsquo;s the Jobs page — the new table follows the same visual style:\nWhat the Table Shows Column What It Shows Job ID Link to the job detail page Description Stage name and description Submitted Submission timestamp (sortable) Duration Human-readable duration (sortable) Stages: Succeeded/Total Stage completion with failed/skipped counts Tasks: Succeeded/Total Progress bar with task-level breakdown The table header shows \u0026ldquo;Associated Jobs (N)\u0026rdquo; so you immediately know how many jobs the query spawned. Click the header to collapse or expand the section — the state persists across page reloads via localStorage.\nColumns are sortable. Click Duration to find the slowest job. Click Submitted to see execution order. The Stages and Tasks columns use the same progress bar style as the main Jobs page, keeping the visual language consistent.\nWhy This Matters The SQL execution detail page is where engineers go when a query is slow or failing. The questions are always the same:\nHow many jobs did this query create? — Now visible in the section header. Which job is the bottleneck? — Sort by Duration. Are jobs still running? How far along? — Task progress bars show this instantly. Did any stages fail? — The Stages column shows failed/skipped counts inline. What\u0026rsquo;s each job actually doing? — The Description column shows stage names. Previously, answering any of these required navigating away from the execution page. Now they\u0026rsquo;re all answered in one table, on the same page, without a single click.\nBonus: Concise Progress Bar Labels The same PR also fixes a long-standing readability issue with progress bars across the entire Web UI. When tasks are killed, Spark used to display the full kill reason — including stack traces — inside the progress bar label:\n[====\u0026gt; ] 45/100 (5 killed: org.apache.spark.SparkException: Job 3 cancelled because SparkContext was shut down at org.apache.spark.scheduler.DAGScheduler...) Now it shows a concise label with the detail available on hover:\n[====\u0026gt; ] 45/100 (5 killed) ↑ hover for full reason This applies to progress bars on the Jobs page, Stages page, and the new SQL execution jobs table. The truncated reason (up to 120 characters) is kept in the tooltip.\nPart of a Bigger Modernization This is one piece of the broader SPARK-55760 Web UI modernization. Other recent improvements include:\nDark mode — one-click toggle, OS preference detection (SPARK-55766) Compact SQL plan visualization — edge row counts, clickable metric panels (SPARK-55785) Offcanvas detail panels — slide-out executor views (SPARK-55767) Bootstrap 5 Collapse API — replacing custom JS collapse across all pages (SPARK-55773) Bootstrap 5 migration — from 4.6.2 to 5.3.8 (SPARK-55761) The goal is simple: the Spark Web UI should be as good at helping you understand your queries as Spark is at running them.\nTry It Out This feature is already merged to master and will be available in the next Apache Spark release. If you\u0026rsquo;re building from source, try it today.\nContributed as SPARK-55971 (PR #54768). Feedback and contributions to the Spark Web UI modernization are welcome at SPARK-55760.\n","permalink":"https://yaooqinn.github.io/posts/spark/sql-execution-page-modernization/","summary":"The SQL execution detail page in Spark\u0026rsquo;s Web UI used to show jobs as comma-separated IDs. Now it has a full Associated Jobs table with status, duration, stage progress, and task progress bars — so you can debug SQL queries without clicking through each job individually.","title":"The SQL Execution Detail Page Finally Shows You What Your Jobs Are Doing"},{"content":"If you\u0026rsquo;ve ever stared at the Spark Web UI at 2 a.m. debugging a failing job, you know the feeling: a wall of bright white light hitting your eyes while you scroll through stages, executors, and SQL plans. That era is over.\nApache Spark\u0026rsquo;s Web UI now supports dark mode, landing in the master branch as part of SPARK-55766. Toggle it with a single click. It remembers your preference. It respects your OS settings. And it covers every page — Jobs, Stages, SQL, Executors, Environment, and beyond.\nWhy Dark Mode? Dark mode isn\u0026rsquo;t a cosmetic trend — it\u0026rsquo;s a developer productivity feature. Here\u0026rsquo;s why it matters for Spark:\n1. Reduce Eye Strain During Long Debug Sessions Spark jobs can run for hours. When something goes wrong, engineers spend extended periods navigating the Web UI — examining DAG visualizations, reading executor logs, and tracing SQL query plans. A dark interface reduces the contrast between the screen and a dimly lit room, which is exactly where most late-night debugging happens.\n2. Respect Developer Preferences Modern developer tools have universally adopted dark mode — VS Code, GitHub, IntelliJ IDEA, terminal emulators, even operating systems. When a developer has their entire environment in dark mode and then opens the Spark UI in a browser, the bright white page is jarring. The Spark UI should feel like a natural part of the developer\u0026rsquo;s workflow, not an exception to it.\n3. It\u0026rsquo;s What the Community Has Been Asking For Dark mode has been one of the most requested UI features in the Spark community. As the Web UI moves toward Bootstrap 5 modernization (SPARK-55760), the infrastructure finally exists to support it properly — without hacks, without custom CSS themes, and without maintenance burden.\nWhat It Looks Like Here\u0026rsquo;s the Spark Web UI in both modes, side by side:\nJobs Page Light Dark SQL Query Plan Visualization Light Dark Executors Page Light Dark Environment Page Light Dark How It Works — The Philosophy We deliberately chose the simplest, most maintainable approach:\nLeverages Bootstrap 5 color modes. No custom CSS color system. No separate stylesheet. Bootstrap\u0026rsquo;s data-bs-theme attribute handles 95% of the work — buttons, cards, tables, navbars, and text all adapt automatically.\nRespects system preferences. On first visit, the UI checks prefers-color-scheme — if your OS is in dark mode, Spark follows suit. No configuration needed.\nRemembers your choice. Click the toggle once, and localStorage persists your preference across every page and session.\nNo flash of unstyled content (FOUC). An inline script runs before the page renders, so you never see a white flash when loading in dark mode.\nThis philosophy mirrors how the best developer tools handle theming:\nGitHub introduced dark mode in December 2020 with a similar approach — respecting system preferences, persisting choice, and using CSS custom properties rather than a parallel stylesheet. VS Code has had dark mode since day one, treating it as a first-class feature rather than an afterthought. Grafana, another tool engineers stare at for hours, defaults to dark mode entirely. The lesson from all of these: dark mode isn\u0026rsquo;t optional for developer-facing tools. It\u0026rsquo;s expected.\nPart of a Bigger Modernization Dark mode is one piece of a broader effort to modernize the Spark Web UI under SPARK-55760. Other improvements landing alongside it include:\nCompact SQL plan visualization with a clickable detail side panel for metrics Offcanvas panels for executor detail views Table hover effects for better row readability Bootstrap 5 utility classes replacing legacy CSS A global footer showing version, uptime, and user across all pages The goal is straightforward: the Spark Web UI should feel as modern as the engine it represents. Spark processes petabytes of data with cutting-edge distributed computing — its UI should reflect that quality.\nTry It Out Dark mode will be available in the next Apache Spark release. If you\u0026rsquo;re building from source, it\u0026rsquo;s already on the master branch.\nToggle it by clicking the ◑ button in the navbar. That\u0026rsquo;s it.\nThis feature was contributed as part of SPARK-55766. Feedback and contributions to the Spark Web UI modernization are welcome at SPARK-55760.\n","permalink":"https://yaooqinn.github.io/posts/spark/dark-mode-spark-ui/","summary":"Apache Spark\u0026rsquo;s Web UI now supports dark mode — a long-awaited quality-of-life improvement for developers who spend hours debugging jobs. Here\u0026rsquo;s why we built it and what it means for the Spark community.","title":"Dark Mode Comes to the Apache Spark Web UI"},{"content":"The SQL tab in the Spark Web UI has always had a query plan visualization. It shows you the physical plan as a DAG — operators as nodes, data flow as edges. In theory, it\u0026rsquo;s one of the most powerful debugging tools in Spark. In practice, it\u0026rsquo;s been nearly unusable for complex queries.\nThat changes now. SPARK-55785 reimagines the SQL plan visualization with compact layouts, interactive metric panels, and edge labels that surface data flow at a glance.\nThe Problem The old visualization crammed every metric into every node label:\nHashAggregate number of output rows: total (min, med, max) 5,000,000 (1,250,000, 1,250,000, 1,250,000) time in aggregation build: total (min, med, max) 2.3s (500ms, 575ms, 650ms) peak memory: total (min, med, max) 512.0 MiB (128.0 MiB, 128.0 MiB, 128.0 MiB) avg hash probe bucket list iters: ... 1.2 (1.1, 1.2, 1.3) Multiply that by 30+ operators in a real-world query, and you get a wall of text where nothing stands out. The most critical information — which operator is slow? where do rows explode? where does the data get filtered down? — is buried in visual noise.\nThe Solution Compact Mode: See the Forest, Not the Trees The new default view shows operator names only. No metrics clutter. The plan structure is immediately readable:\nEach node is just its operator name. Each cluster shows the WholeStageCodegen stage number and total duration. The plan fits on one screen instead of requiring endless scrolling.\nEdge Labels: Data Flow at a Glance The most impactful addition isn\u0026rsquo;t in the nodes — it\u0026rsquo;s on the edges. Row counts now appear on every edge, showing exactly how much data flows between operators:\nThis makes several classes of performance problems immediately visible:\nJoin explosions — 1M × 500K inputs producing 5B rows? You\u0026rsquo;ll see it instantly on the edge label. Filter effectiveness — Does your filter reduce 5B rows to 400? The edge tells you without clicking anything. Aggregation impact — See exactly how much your GROUP BY compresses the dataset. These are the questions engineers ask first when debugging slow queries. Previously, you had to mentally trace through metric tables or use EXPLAIN output. Now the answer is right there on the graph.\nClick for Details: Metric Side Panel Need the full picture? Click any node. A side panel slides in with structured metric tables:\nThe panel shows:\nTotal / Min / Med / Max breakdowns per metric Operator description on hover for disambiguation (useful when you have multiple HashAggregate nodes) Cluster click to see all child operator metrics grouped together This design was inspired by Databricks\u0026rsquo; query profiler, adapted for the open source Spark UI.\nToggle: Your Choice Prefer the old detailed view? A checkbox lets you switch between compact and detailed modes. In detailed mode, metrics are rendered inside the graph nodes (at 10px font to keep things readable):\nYour preference is saved in localStorage — the UI remembers how you like it.\nDesign Decisions A few choices worth calling out:\nWhy edge labels over node metrics? Because the most important performance signal is data volume between operators, not the internals of a single operator. Edge labels answer \u0026ldquo;what happened to my data?\u0026rdquo; at the plan level. Node metrics answer \u0026ldquo;why is this specific operator slow?\u0026rdquo; — and those belong in the detail panel, not the overview.\nWhy plain text labels instead of HTML? The old visualization used dagre-d3\u0026rsquo;s labelType: \u0026quot;html\u0026quot; for rich formatting inside nodes. This caused rendering inconsistencies, made dark mode support harder, and produced oversized nodes. Plain text labels are lighter, more predictable, and let dagre-d3 auto-size nodes correctly.\nWhy a side panel instead of tooltips? Metric tables can be large — 6+ metrics with Total/Min/Med/Max columns. Tooltips would clip or overlap. A fixed side panel provides stable, scrollable space and doesn\u0026rsquo;t obscure the graph.\nWhat\u0026rsquo;s Next This is part of the broader Spark Web UI Modernization effort. Future improvements under discussion include:\nShowing data flow paths for a single selected operator Highlighting bottleneck operators by color (time/rows heatmaps) Edge annotations for data size in bytes, not just row counts This feature was contributed as SPARK-55785 (PR #54565). Thanks to @sarutak and @gengliangwang for the review and the suggestion that led to edge row count labels.\n","permalink":"https://yaooqinn.github.io/posts/spark/sql-plan-visualization/","summary":"The Spark SQL plan visualization just got a major upgrade — compact node labels, clickable metric panels, and edge row counts that make join explosions immediately visible.","title":"Rethinking SQL Plan Visualization in Apache Spark"},{"content":"👋 Hi, I\u0026rsquo;m Kent Yao I\u0026rsquo;m an open source enthusiast and a passionate contributor to the Apache Software Foundation ecosystem. I focus on big data, distributed SQL engines, and building open source communities.\nRoles 🧑‍🤝‍🧑 ASF Member 🍼 Apache Incubator PMC Member 🦊 Apache Kyuubi PMC Chair, Vice President ✨ Apache Spark PMC Member 🚢 Apache Submarine Committer 🧱 Databricks Beacons Program Member Open Source Journey Year/Month Organization Event 2024/10 Apache Software Foundation Apache Cloudberry Mentor and PPMC Member 2024/08 Apache Software Foundation Apache Spark PMC Member 2024/07 Apache Software Foundation Apache Polaris PMC Member 2024/06 Databricks 2024 Databricks Beacons 2024/03 Apache Software Foundation Apache Amoro(Incubating) Mentor and PPMC Member 2024/03 Apache Software Foundation ASF Member (news) 2024/01 SegmentFault 思否 / 开源社 2023 中国开源先锋 33 人 2024/01 Apache Software Foundation Apache Gluten PMC Member 2023/12 OpenAtom Foundation 2023生态开源项目: Apache Kyuubi 2023/11 Apache Software Foundation Apache Incubator PMC Member 2023/10 NetEase Corp. NetEase Technology Award 2023/09 中国信息通信研究院 2023 OSCAR 尖峰开源人物 2023/05 中央网信办信息化发展局 2022年中国开源创新大赛 - 二等奖 2022/12 Apache Software Foundation Apache Kyuubi PMC Chair, Vice President 2022/12 Apache Software Foundation Apache Kyuubi becomes ASF Top-Level Project 2022/10 NetEase Corp. NetEase Technology Award 2022/09 中国信息通信研究院 / 中国通信标准化协会 2022 OSCAR 开源产业大会尖峰开源项目: Apache Kyuubi(Incubating) 2022/06 ACM SIGMOD The 2022 ACM SIGMOD Systems Award 2022/05 中国信息通信研究院 可信开源社区共同体: Apache Kyuubi(Incubating) 2022/02 中国科学技术协会 2021 \u0026ldquo;科创中国\u0026quot;开源创新榜: Apache Kyuubi(Incubating) 2021/10 NetEase Corp. NetEase Technology Award 2021/08 Databricks Databricks Beacons Program Member 2021/06 Apache Software Foundation Apache Kyuubi(Incubating) PPMC Member 2021/06 Apache Software Foundation Donated NetEase/Kyuubi into the Apache Incubator 2021/02 Apache Software Foundation Apache Spark Committer 2020/12 Apache Software Foundation Apache Submarine Committer ","permalink":"https://yaooqinn.github.io/about/","summary":"About Kent Yao","title":"About"},{"content":"🦊 Apache Kyuubi Role: PMC Chair, Vice President\nA distributed and multi-tenant gateway to provide serverless SQL on data warehouses and lakehouses. I donated the original project from NetEase into the Apache Incubator in 2021, and it graduated as an ASF Top-Level Project in 2022.\n🔗 GitHub · Website\n✨ Apache Spark Role: PMC Member\nA unified analytics engine for large-scale data processing. I\u0026rsquo;ve been contributing to Spark SQL, focusing on SQL compatibility, data types, and configuration improvements.\n🔗 GitHub · Website\n🌟 Apache Gluten Role: PMC Member\nA plugin to double SparkSQL\u0026rsquo;s performance by offloading the SQL engine\u0026rsquo;s execution to native engines. Gluten graduated as an ASF Top-Level Project.\n🔗 GitHub · Website\n🧊 Apache Polaris Role: PMC Member\nAn open source catalog for Apache Iceberg. Polaris graduated as an ASF Top-Level Project.\n🔗 GitHub · Website\n🫐 Apache Cloudberry Role: Mentor, PPMC Member\nA next-generation unified database for analytics and AI, forked from Greenplum Database.\n🔗 GitHub · Website\n🐻‍❄️ Apache Amoro Role: Mentor, PPMC Member\nA lakehouse management system built on open data lake formats.\n🔗 GitHub · Website\n🚢 Apache Submarine Role: Committer\nA unified AI platform that allows engineers and data scientists to run machine learning and deep learning workloads.\n🔗 GitHub · Website\n","permalink":"https://yaooqinn.github.io/projects/","summary":"Open source projects","title":"Projects"}]