SpecsModelExploreRequirements
SpecsModelExploreRequirements
Spark Runner
Execution
Notes
Spark Runner
Overview
Guide
Twin streams
Links and Communications

Spark Runner

The Spark Runner is the execution engine of the batch side. It takes a job dispatched by the Job Scheduler, submits a Spark application, and supervises it until the curated output is handed to the Export Writer.

Elastic by design

Executors scale with the size of the day being processed. A normal day uses a modest pool; a backfill of weeks spins up far more parallelism and releases it when the job completes.

Execution

The runner builds the application config — input partitions, executor memory, parallelism — and submits it to the cluster.

Notes

  • Reads raw telemetry for the target date and produces curated, deduplicated records with late data folded in.
  • Resource limits are bounded so a large backfill cannot starve scheduled daily runs.