Skip to main content

Engine architecture

Quanton combines native execution with an understanding of how your tables are stored. It uses that knowledge to reduce the data a Spark job must read, process, and write, then executes the remaining work efficiently through fast relational operators and optimized SQL plans. Your Spark application code remains unchanged.

Storage-aware execution​

Spark performance tuning often starts and ends with the job and its stages: adjust partition counts, choose a join strategy, reduce skew, or allocate more executors. Native acceleration adds another improvement by executing relational operators in efficient C++ or Rust code. These techniques reduce execution time, but storage access and shuffling data can still dominate the job.

Data warehouses gain efficiency by optimizing execution and storage together. File layout, column statistics, indexes, and the query engine work together to avoid unnecessary reads and processing. Storage knowledge directly reduces the work the engine must perform.

Quanton brings a storage-aware approach to Spark performance on open lakehouse tables. It uses open table format metadata, table layouts, and available indexes to select data, reduce join input, and limit data movement. Native readers, relational operators, and writers then process the remaining data in columnar form. Planning, storage access, and execution become parts of the same optimization problem.

Consider a MERGE that updates a small set of records in a large table. Faster operators help process the target scan, but a suitable index can identify the files and row groups that contain matching records. Quanton combines that smaller scan with a columnar merge path that reduces shuffle. The benefit comes from both doing less work and using fewer resources for the work that remains.

Optimizing the full Spark pipeline​

Quanton optimizes the full pipeline across all stages, instead of narrowly focussing on transformations alone. These optimizations preserve Spark's stage dependencies: a dependent stage waits for the required shuffle outputs from its upstream stages.

The diagram groups optimizations by type of work. Extract: read storage with I/O pipelining and less scheduling overhead. Transform: reshape plans, execute native operators, use indexes, and handle resource pressure. Load: append or merge data with columnar compaction and less shuffle. Table metadata, statistics, and indexes support these optimizations. The groups do not represent concurrent Spark stages.

Native C++ or Rust operators accelerate operator execution. Quanton also optimizes reads, query plans, and writes. The diagram groups these optimizations by type of work. Open full-size diagram ↗

Extract: read from storage​

Reading a lakehouse table involves remote storage requests, decompression, decoding, and table-format operations such as applying deletes or merging log files or reconciling deletion vectors. These operations use different resources. Storage requests spend time waiting on the network, while decoding consumes CPU. If a reader waits for each request before continuing, CPU capacity can sit idle even when there is more work to do.

Quanton uses I/O pipelining to overlap storage requests with processing inside the read path. While the reader decodes available data, requests for subsequent input can remain in flight. This keeps data flowing to the CPU and makes better use of available network bandwidth.

The design draws on principles from staged event-driven architecture (SEDA): separate operations with different resource needs and manage their concurrency. Quanton applies this principle to storage access and data processing, so remote I/O does not force the entire reader to wait. This pipelining operates within Spark's stage dependencies.

Two other optimizations support this read path:

  • Native reads: Quanton's reader decodes data in columnar batches and applies table-format operations. This reduces CPU overhead and supplies data in the form used by native operators.
  • Scan scheduling: Quanton distributes scan work to reduce scheduling overhead and delays from uneven file sizes. This helps prevent a few slow tasks from holding up the rest of a scan.

Together, these changes help a scan sustain throughput across storage access and decoding. See Lakehouse reads for the implementation approach and a measured scan example.

Transform: execute relational operators​

Quanton combines query plan changes with native vectorized operators. It builds on Velox with rewritten relational operators that improve execution efficiency and memory management, plus optimizations for cloud CPU architectures.

SQL plan reshaping​

Quanton applies rule-based optimizations that adapt to table metadata before native operators execute the SQL plan. These rules use information about table layouts, column statistics, and available indexes to reduce repeated work and choose how to access data. Eligible rewrites remove duplicate scans, collapse self-joins, and reduce unnecessary shuffles while preserving query results.

For example, metadata can show whether partition pruning will narrow a join's target scan enough, or whether an available index offers a more selective read. Quanton adapts the plan to the table and query, so less data reaches downstream operators.

Your SQL stays the same. Native operators execute the reshaped plan, combining fewer reads and shuffles with faster processing of the remaining data.

Native rollup​

A ROLLUP calculates subtotals and a grand total, such as sales by product, category, and overall. Spark's Expand operator copies each input row for every grouping level; an eight-key rollup produces nine copies before aggregation, adding CPU and memory work.

Quanton's native rollup operator builds higher-level totals from results it has already aggregated. This avoids multiplying the original rows and reduces the work needed to produce the same report. See Native rollup for an example.

Index-aware joins​

Joining a small batch of keys against a large table can require scanning much of that table just to find matching rows. Quanton adds index-aware join, a new relational operator beyond the join operators available in standard Spark.

The operator looks up source keys in a suitable table index, which identifies the target files and row positions. It then reads the row groups containing matches, reducing the data scanned and passed into the join. This is especially useful for incremental updates whose keys are scattered across a large table. See Index-aware joins and MERGE for the lookup sequence.

Memory pressure​

Large joins and aggregations can push native memory beyond an executor's limit, causing it to be killed and tasks to restart. Quanton tracks native allocations against an executor memory budget and uses an allocator designed to reduce fragmentation and retained memory. Operators spill data to disk under pressure, helping the job continue with extra I/O rather than fail; see Native memory and spill for details.

Load: append, merge, and compact data​

A daily batch of updates may change only a small fraction of a table, yet a conventional MERGE can shuffle large amounts of target data. The network cost can dominate the job even when there is little new data to write.

Quanton combines index lookups to locate affected records with a low-shuffle columnar MERGE path to apply changes. The index reduces the search work, while the merge operator reduces target data movement across the cluster. Affected files still need the writes required by the table format, but locating and updating records takes less work. See Low-shuffle MERGE for a worked example.

Quanton can build indexes on the fly from explicit user hints or by automatically detecting write patterns and identifying merge columns. It maintains these indexes as data changes, keeping them consistent with the table. These indexes accelerate record lookups for Iceberg MERGE INTO and DELETE statements.

For appends, native writers keep data columnar through the write path. For compaction, Quanton combines base data and accumulated updates in columnar batches, reducing the work later reads must repeat. The writers maintain the data files and metadata required by Iceberg and Hudi.

Evaluate your workload​

The benefit depends on the job's main costs. A large scan, an aggregation, and an incremental merge use different optimizations.

Review the benchmark results for measured comparisons and configuration details. Follow Benchmarking to run the TPC-DS suite. Use the Local Quickstart to set up Quanton.