← All parts

Capacity Planning

Design a Distributed Job Scheduler · Part 5
Problem context

Design a system that runs jobs at a chosen time: "run job X at time T" once, or on a recurring schedule like "every day at 09:00." The work must run reliably at scale, even when the machines running it crash, and it must run effectively once — never silently skipped, never harmfully duplicated.

In scope: submitting one-off and recurring jobs, dispatching due jobs to workers, crash recovery, and effectively-once execution. Out of scope: the workers' own job logic, and dependencies between jobs (a workflow/DAG engine, covered briefly as a variant).

What a strong answer sounds like

State the decision, connect it to a requirement, and name the tradeoff. Keep the design focused on the workload in the prompt.

Ready for the capacity planning interview?

The AI interviewer asks about this part of Design a Distributed Job Scheduler. The interviewer guides you through topics one by one.