Engineering note 3 · Operations

A check that never pinged never alerts

Moving scheduled trading jobs from a laptop to a server, and the ways a system can report that everything is fine while it is quietly doing the wrong thing.

October 2026 · 3 min read

For its first months the system ran on a laptop: Windows Task Scheduler started the jobs, and a supervisor process caught any it missed. It worked most of the time. The failures it did have are worth writing down, because none of them looked like failures when they happened.

On wake, everything at once

When a Windows machine wakes after missing scheduled tasks, it does not run them in order. It runs all of them in the same second. After one long gap, every job logged an identical start time, to the second. The scheduler's health report looked perfect: every job had run.

The problem was the simultaneity. The daily trading run and the weekly price refresh both read and write the same price cache, and running together they lost eleven writes to file-access errors. Nothing in a run-or-did-not-run check can see that.

Two launchers, 250 milliseconds apart

The daily run had two independent ways to start: the scheduled task, and the supervisor's catch-up for days the machine had been asleep. One day they fired 250 milliseconds apart, and the run failed.

The fix was a lock. The pipeline now takes an exclusive lock for its whole duration. A second copy that finds it taken exits with a distinct code, and the wrapper records that as skipped, not failed, so a harmless collision does not break the reliability record. Anything else that touches the shared data takes the same lock.

Three missing days

A laptop is sometimes simply off. One week in September, three trading sessions passed with no run at all. Nothing alerted until the next run, because nothing was running to raise an alert. The strategy monitor found the gap afterwards and now says so in every report: the drawdown and exposure figures cover an incomplete series, and can read no worse than the gap allows. That is honest, but it is a report, not an alarm.

The jobs have since moved to an always-on server, where a machine being asleep is no longer a reason to miss a day.

Three server rules that fail silently

Scheduled jobs on Linux are written as cron entries. Each of these mistakes produces a job that runs and appears to succeed:

  1. Call the project's Python by its full path. A bare python finds the system interpreter, which lacks the database driver. The code then falls back to writing JSON files, and the run reports success.
  2. Set the server's clock to the market's time zone first. Cron fires on the system time zone. A TZ= line in the schedule changes what the jobs see, not when they start, so on a UTC server a 16:45 job runs at 22:15 Indian time.
  3. Make the broker check actually renew. A "validate only" mode checks the broker session and never logs in again, so the first expired session stays expired.

The alarm that could not ring

Every job now reports its exit code to healthchecks.io, an independent service that emails when a job is late, fails or never reports. It is the outside alarm: it works even if the whole server is down.

It has one property that is easy to miss. A check that has never received a ping does not alert. It sits in a "new" state indefinitely. If a job's very first scheduled run never happens, its check never goes red. On the day of the move, the new checks sat in that state for hours, looking armed when they were not.

The fix is one line per check: send a single manual ping the moment the check is created. From then on, silence is an alarm.

Backups that are actually tested

The database is backed up every night, and once a month a scheduled drill restores the latest backup into a scratch database and compares it with the real one. On the very first day, the drill caught a backup that had silently failed to upload. That is the case the drill exists for, and without it the gap would have shown up only when the backup was needed.

The lesson

Every alarm needs proof that it can fire, and every "success" needs proof that the right thing succeeded. A green dashboard tells you nothing about either.