Spark versus Luna Evaluation, July 2026

Question

For one fixed, read-only corpus task, did a latency-focused Spark profile at high reasoning effort outperform a low-effort Luna profile enough to justify using it by default?

Result

Both runs scored 40 of 40 deterministic answer atoms and passed the required schema. The Spark run completed in 24.89 seconds and the Luna run in 56.12 seconds, making Spark 2.25 times faster in this sample.

The faster run used 15,277 output tokens, including 13,729 reasoning tokens, and made 13 successful shell calls. The lower-cost run used 2,585 output tokens, including 664 reasoning tokens, and made 3 successful shell calls.

Interpretation

This sample supports Luna low as the economical default for clear, repeatable, batchable read-only work. Spark high effort may be useful when one tiny lookup blocks an interactive workflow and latency matters more than usage. The sample does not show an accuracy advantage for Spark and does not isolate the value of high reasoning effort because no lower-effort Spark control was run.

Do not generalize this result to writing, architecture, policy, security, or other risk-bearing work.

Limitations and Retest

This was one task with one run per profile. Two fixture cardinality defects and a nested instruction-scope difference reduced comparability. Retest with at least three frozen prompts, corrected fixtures, and a lower-effort Spark control when the CLI, model, or profile changes or before granting write scope.

Citations