Engineering notes · October 2026 · engine v26

How Aevox reads a metabolic test

Turning messy device exports into numbers a coach can act on. Deterministic code does the math, a model proposes, and physics checks and a physiologist decide.

The problem

A cardiopulmonary exercise test (CPET) is a ramp to exhaustion with breath-by-breath gas analysis. From one table of breaths, a lab wants VO₂max, the two ventilatory thresholds that set training zones, how the athlete burns fat and carbohydrate, and a retest plan. Aevox reads the file the lab’s metabolic cart exports and turns it into that report.

The files are hostile. The table sits somewhere different for every vendor. Three columns can all be called “VO2”. A column labelled in minutes can hold Excel day-fractions. Some analysers have no CO₂ channel at all. Classic threshold algorithms return confident nonsense on bad input: my first O₂-only file produced a high-confidence threshold from a column of zeros. So most of the code is about knowing when a number is wrong, and saying why.

Aevox is deployed in five labs on five cart vendors and is in production at Hybl Performance Center, part of CommonSpirit Health and powered by EXOS. I built it, and I run every deployment: the first call, the lab’s files as they actually arrive, and the fix.

The shape

Aevox upload pipeline A cart export goes through ingest, the engine, physiologist confirmation and the report. When ingest refuses a sheet, a model proposes a column map that must pass a physics gate. The model also reads the thresholds inside bounds before the physiologist confirms. Cart export, as uploaded Ingest table · columns · time · units Engine filters · VO₂max · ensemble Physiologist confirms Report measured · provisional · withheld Model proposes a column map Physics gate ranges · HR · RER identity Model reads thresholds one VO₂ value, in bounds refused pass fail → refusal with a reason
Solid lines are deterministic code running in the clinician’s browser. Dashed lines are model calls through a thin API. Every dashed line ends at a check or a person.

The analysis engine is plain TypeScript running in the clinician’s browser, inside a React app. A thin serverless API handles auth, storage, sharing and model calls (Vercel, Clerk, Neon Postgres). There are 900+ tests, and the engine has gone from v2 to v26 since July.

Life of an upload

  1. Read every file safely. Each file gets its own try/catch, so a crash becomes a “skipped” note instead of a blank page. A PDF renamed .xlsx is caught by its magic bytes.
  2. Find the table and name the columns. The table is found by numeric density. Headers are scored onto 31 fields by name and unit together, units convert on read, and the time scale is measured from the data.
  3. If the rules refuse, maybe ask the model. When a sheet is refused for mostly unrecognised headers, Claude gets each column’s name, unit, five samples and summary statistics, never the file or the athlete, and proposes a map. The map applies only if the full columns pass physics checks, and AI-assigned columns are marked for verification.
  4. Analyse. Outlier filtering, the cart’s own smoothing filter, a 30-second VO₂max window per bout, and an ensemble of threshold detectors (five for each threshold on full-gas files, fewer when there is no CO₂ channel).
  5. Second read. The model returns one VO₂ value per threshold. Code rejects anything outside physiological bounds and falls back to the engine with a banner.
  6. Confirm. The physiologist sees the raw curves, every candidate, the model’s reasoning and sliders. Nothing renders before they confirm.
  7. Decide what each card may claim. A pure function sets each of 19 report cards to measured, provisional or withheld, with the reason on page one.
  8. Save, share, review. The report, the original bytes and the engine version are saved together. The athlete gets a signed link. A separate model review reads the new report once, without the athlete’s name, and never blocks it.

Ten decisions

1. Run the engine in the browser; keep the server thin.

WhyNo compute bill, one render path for the clinician, the athlete and the PDF, and a keep-local mode becomes possible.

TradeoffLarge payloads, and the server stores what the browser computed. Integrity comes from archiving the original file and being able to re-run it.

2. Read files by shape and units, not vendor templates.

A column scores +0.6 for a name hit and +0.4 for the unit, loses 0.35 for a contradicting unit, and is accepted at 0.5. One vendor writes three different columns all headed “VO2”; only the units tell them apart.

TradeoffThe weights are hand-set, and the matcher is greedy: a column that loses a collision isn’t tried for its next-best field.

3. Measure the time scale; don’t trust the label.

A breath lasts 60/RR seconds, so seconds per unit = mean(60/RR) ÷ mean(Δtime), snapped to 1, 60, 1440, 3600 or 86400. One major cart labels its time column in minutes but writes Excel day-fractions meant as MM:SS. Before this, that file’s time axis came out 2.5× too long, which silently broke every rolling window, VO₂max included.

TradeoffIt needs a breathing-rate column; without one, it falls back to number format, then label, then a flagged guess.

4. The model proposes; physics decides.

A proposed column map applies only if the values fall in human ranges, heart rate tracks VO₂ (r ≥ 0.5), and, for whichever channels are mapped, the cart’s own identities hold: RER = VCO₂/VO₂, constant body mass, VE = RR × VT. The model’s self-reported confidence is ignored. A mislabelled column can’t satisfy RER = VCO₂/VO₂ over hundreds of breaths; on real files the identity closes to about 0.003, against a warning at 0.03. The worst-case mapper call costs about fifteen cents.

What still fools itAnything proportional to the real channel. A %VO₂max column mapped as VO₂/kg tracks heart rate, and the implied body mass comes out a steady 35 kg, inside the plausible band. The fix is to check implied mass against a known mass, and to cap confidence when nothing ties VO₂ to an independent channel.

5. Thresholds: an ensemble, a bounded model read, then a person.

Each detector fails differently, so out-of-band candidates are dropped and the rest form a weighted median. The model returns one VO₂ number, never a time, heart rate or power, and code snaps it, or any slider, to a real breath.

TradeoffOne more step per test. In Aevox Labs, the self-serve version, the engine sets the defaults and the model’s read is a flagged second opinion; I’m bringing that to every deployment.

6. Withhold what a missing input makes false; caveat what a suspect input makes uncertain.

A percentile of a non-maximal peak is false even with a caveat, so it’s withheld, and the card says what input would unlock it. Rules go through one setter that can only move a card toward more doubt, and states are derived at render, never stored, so nobody can hand-set “measured”.

7. Version every number and keep the bytes.

Any change that moves a number bumps the engine version, which is saved with the original file. Re-derivation is read-only and flags headline values that moved more than 1%, so an engine change can’t silently rewrite a signed report. Every bump records a before/after diff across 32 real test sessions: a sign fix in one threshold method left 26 of them unchanged and moved VT1 on the other six, so the diff shows exactly which athletes a fix touches.

8. The model review never blocks, and its findings become rules.

A coach never waits on a verdict. The review separates engine defects from data limits. On two trial labs’ files it found three missing rules (an unknown incline, no belt speed, detectors that disagree), which are now code pinned by tests.

TradeoffErrors are caught later, by me. In Aevox Labs only the riskiest uploads are held, and there the review can only downgrade.

9. Keep-local, enforced on both sides.

Some clinics can’t send test data to a vendor. One switch stops every save, link, lookup and model call: browser helpers refuse and API routes answer 403, so even a cached page is refused. A test fails if any file that makes a network request hasn’t been classified against the switch.

10. One codebase; a lab is configuration, never a fork.

An early branded fork fell 19 engine versions behind, so new labs build from one core, each with its own isolated deployment, auth and database provisioned by one script. One push upgrades all of them.

TradeoffA bad push also reaches all of them, which makes the deploy check the most important thing in CI.

From the field

  • A unit that was 60× off. One export labelled energy in kcal per hour, the mapper filed it as kcal per minute, and fat oxidation came out 60× too high. Units now convert on read, and a device energy column more than 15% off the Weir equation is dropped.
  • A new cart, same day. A lab’s ParvoMedics export was refused because its gas-state row (STPD/BTPS) was read as the column names. Once read, heart rate typed only at stage ends would have dropped 87 of 96 rows. The fix shipped as an engine bump with a regression test the day the file arrived.
  • A product card from a drawing. A lab sketched the lactate comparison they did by hand; it became a report card the same week.
  • A deploy that lied. I hit the hosting plan’s function cap, and production kept serving a three-week-old build while two merged commits failed to deploy. My probe passed because the old build still answered 200. I folded handlers behind one router per group, URLs unchanged. The lesson: a deploy check compares the serving commit, not the status code.

How I build with agents

Much of the code is written by Claude Code agents, and the commits say so. Aevox Labs, the self-serve multi-tenant version, was built by ten parallel agent lanes working from a written plan; I owned the plan, the physiology review and the QA rounds. Its upload agent reads unknown files through 13 fixed engine operations, and a confirmed recipe becomes that lab’s adapter. What keeps it honest is the same machinery as the engine: cross-fixture diffs on every bump, a harness of synthetic bad reports, and a person who can read the test. The company runs the same way: scheduled agents draft and watch, and nothing reaches a customer without my sign-off.

What I’d do next

  • Threshold ground truth. Confirmation is mandatory because there is no labelled reference. Pre-filled sliders anchor the clinician, so only moved values are usable labels. The next step is a blinded, hand-labelled set; one detector lands in band on only 7 of 26 files, so its double weight goes when labels arrive.
  • A mapper eval for free. Scramble the headers on files the rules already read, then score accuracy, false accepts and false declines, and gate every model upgrade on it.
  • The upload flow as a state machine. One component holds the pipeline, and its step order caused three bugs. A step reducer with tested transitions replaces it.
  • Append-only reports. Every save becomes a new version, and any edit clears sign-off.
  • Count every fallback. Each deterministic floor should increment a counter and alert, so a silent fallback can’t hide a broken model call.

On a new team with a problem like this

I’d get ten real files and the customer’s own printouts and build the ground-truth comparison before any features. Then find the identities that let the data check itself, ship refusals with reasons before guesses, and version every number from day one.

Questions or a file your software can’t read: jackmis610@gmail.com.