Health insights that wait for evidence.
Quorum reads your health data and looks for what's actually there. Axiom, its agent, reasons from first principles — and tells you plainly when the answer is that it doesn't know yet.
The problem
Most health apps are built to alarm you.
An elevated reading on a single night is noise. Sold back to you as a notification, it becomes anxiety — and after enough false alarms, you stop reading the ones that matter.
This isn't a matter of taste. Alert fatigue is a documented clinical harm: as the volume of alerts rises, clinicians override more of them, including the ones that were right. Ancker et al. measured it in 2017, and the finding generalizes uncomfortably well to a wrist that buzzes every morning with a verdict on your night.
The competition sells urgency. Quorum sells restraint. Every design decision below follows from that one, including the ones that make the product look less capable than it is.
Quorum sensing
The name is the thesis.
Bacteria don't act alone. Each cell releases a signal molecule and listens for how much is coming back. Below a threshold, nothing happens. Once enough cells agree — a quorum — the whole population acts at once.
It is one of biology's most elegant answers to a problem every health app has and almost none solve: how do you know when a signal is real enough to act on? A single cell can't tell. Neither can a single night's heart rate variability.
So Quorum does what the cells do. It accumulates. A finding stays quiet until it has crossed a threshold that was set before the data arrived — enough days, enough persistence, enough effect size to clear the noise. Restraint isn't the absence of an opinion. It's the discipline of earning one.
Axiom
An agent that shows its work.
Axiom doesn't hand you a score. It investigates — and streams the reasoning live, so you can watch it check, connect, rule out, and occasionally give up.
An axiom is a starting truth: the thing you reason from. That's the job. Every session follows the same five-step protocol, and the interesting one is step four.
-
Scan
Load the findings. Identify what's most clinically significant, not what's most recent or most alarming.
-
Investigate
Don't just report a finding — verify it, contextualize it, and check what else moved with it.
-
Ground
Clinical facts, thresholds, and numbers may come only from retrieved evidence. Never from memory.
-
Ask
When the data shows something it can't explain, Axiom asks you. You have context it doesn't: travel, stress, illness, a new medication.
-
Synthesize
Show what was checked, what was found, what was ruled out — and what remains unexplained.
You
My recovery has been bad all week. What's going on?
Axiom
Your HRV is down 18% against your 30-day baseline, and it started on the 8th. That's a real signal — it's persisted six days, which is past the threshold for noise.
get_baselines · hrv, 30d → mean 62ms, current 51ms
get_correlations · 3 metrics moved together
Resting heart rate rose 4 bpm over the same window and your sleep midpoint shifted 90 minutes later. Those three moving together is a recognizable pattern — but it's consistent with several different causes, and I can't separate them from the data alone.
Did something change around the 8th? Travel, a time-zone shift, illness, or a change in alcohol or training load would each explain this.
Illustrative exchange. Tool names and the streaming protocol are real; the numbers are not from a real user.
What it refuses to do
The constraints are in the code, not the marketing.
Anyone can promise an honest AI. These are the rules Axiom actually runs under — several of them enforced deterministically, outside the model, because a prompt alone isn't a guarantee.
-
It won't cite a paper from memory
Recalling a study without retrieving it is forbidden outright. If Axiom finds itself reaching for a citation to fill a gap in confidence rather than because it actually looked one up, it has to stop and drop the claim.
Never invent clinical claims, paper names, author-year citations, cohort sizes, risk ratios, thresholds, or numeric comparisons from model memory.
-
It won't diagnose you
Findings are observations and associations, never causal claims or conditions. This one isn't left to the model's judgment: a separate layer inspects the output and rewrites diagnostic language before it ever reaches your screen.
Say “your data shows patterns consistent with X,” not “you have X.”
-
It won't dress up a hunch as science
When a pattern is real in your data but has no support in the retrieved literature, Axiom is required to say exactly that — a pattern signal, not a paper-backed one. The two layers are kept structurally separate so “your data shows X” never quietly becomes “research says X.”
-
It won't advertise what doesn't work
Three analysis modules exist in the codebase and are not wired in. They are tagged as such in a machine-readable catalog for the explicit purpose of keeping them out of the interface. Unvalidated scoring weights are labelled design choices, not findings.
-
It won't go quiet on what matters
Repeated low-tier findings are progressively suppressed to prevent the alert fatigue that Ancker's work describes. Urgent and prompt findings are never suppressed, at any repeat count. Quorum gets quieter about noise, never about the things that matter.
-
It won't reach past you
Every tool Axiom has is hard-scoped to the person asking. It cannot access or infer another user's data, and it has no ability to delete data, modify your account, change settings, or send messages on your behalf.
The method
Conventional statistics, held to their actual requirements.
There's nothing exotic here, and that's deliberate. These are well-understood, non-parametric tests chosen because they're robust to the messiness of real wearable data — and, unusually, they're only ever run when their data requirements are genuinely met.
Every sync feeds a seven-phase pipeline: reconcile what changed, rebuild activity-state intervals, derive 210+ features across cardiac, sleep, activity, nutrition, and respiratory domains, then run 12 analyses and 18 composite cross-metric patterns against baselines at 7, 30, 90, and 365 days.
| Question | Method | Won't run without |
|---|---|---|
| Is this unusual for you? | Modified z-score against a 30-day median absolute deviation — robust to outliers in a way a mean isn't. | 14 days + 2-day persistence |
| Is it moving? | Mann-Kendall, a non-parametric trend test that assumes nothing about the shape of the distribution. | 14 days |
| Is it a rhythm? | STL decomposition, separating a repeating cycle from the trend beneath it. | 90 days weekly · 365 annual |
| Do two things move together? | Spearman rank correlation with Benjamini-Hochberg correction — because testing thousands of pairs guarantees false positives without it. | 30 overlapping days · q<0.05 · |ρ|≥0.3 |
| Did something shift? | PELT changepoint detection, locating where a regime actually broke. | Sufficient pre/post window |
What makes a finding worth your attention
Significance isn't a black box. A finding's score is a weighted composite, multiplied by data quality so that a strong signal measured badly can never outrank a modest one measured well — an idea taken from Goldsack's 2020 work on evidence for digital measures.
The largest term asks whether a change is big enough to matter clinically, not merely big enough to be statistically detectable. That distinction — the minimal clinically important difference, after Jaeschke 1989 — is why Quorum stays silent through changes that a p-value would happily call significant. The rest weighs persistence, recency on a 14-day half-life, and agreement across metrics.
An unexpected result
Thirty good days beat eight years.
We assumed history was the moat. It isn't. Almost every analysis saturates within 30 to 180 days — after that, more history buys runtime cost without buying confidence.
What matters is recent and complete. A well-covered month tells Quorum more than a patchy decade, because the tests above care about density and persistence inside a window, not the length of the archive behind it. Annual seasonality is the one genuine exception, and it says so honestly: it will not run until there are a real 365 days.
This is a better product for it. You don't have to have worn something for years to be worth talking to. And confidence is shown the way it's actually earned — “22 of the last 30 days covered,” never a raw sample count engineered to look impressive.
Your last 30 days are well covered.
The evidence base
Every claim traces to a quote.
Roughly two thousand papers, distilled into more than nine thousand individually cited claims: interpretation facts, study findings, device and measurement caveats, and the methodology of the app itself.
The corpus was built by agents, which raises the obvious question. The answer is a rule enforced by a validator, not by good intentions: agents transcribe papers; they never invent findings. Every claim carries a verbatim quote from its source. Missing detail lowers a confidence score rather than getting filled in.
- ~2,000
- papers read and structured
- 9,400+
- individually cited claims
- 328
- supported health data types, of 459 canonical
- 0
- claims accepted without a verbatim quote
We test our own agent for fabrication
An evaluation bank resolves cited DOIs against the live Crossref API and deliberately seeds known failures — a paper with the wrong DOI, a plausible-sounding metric that doesn't exist, an invented tool name — to confirm they get caught. Hallucination is treated as a test case with an expected result, which is the only way to know whether the rules above are actually holding.
Your data
Guarded carefully. Described accurately.
The temptation in health software is to say “bank-grade” and move on. Here is what is actually true, stated at the level of detail that lets you check it.
- The database has no public address. It's reachable only from inside a private network — there is no internet-facing endpoint to attack. Connections are encrypted-only, and credentials live in a secrets manager rather than in code.
- Health data never enters error reporting. Crash and performance telemetry runs with personal information disabled and screenshots off. Authorization headers are stripped before anything is sent.
- Deletion is real deletion. A single endpoint removes your account and its data. Storage is partitioned per user specifically so that erasing one person is a clean operation rather than a search-and-hope.
- Sign-in keys stay on your device. The key material backing your session never leaves the phone, and is excluded from iCloud and device backups.
- The agent is boxed in. Axiom's tools are scoped to you at the infrastructure level, not by instruction. It cannot delete data, change your settings, or message anyone.
- HIPAA-aligned — and we mean the hyphen. Some controls are genuine HIPAA requirements, like automatic logoff under 45 CFR §164.312(a)(2)(iii). Others are defense-in-depth that HIPAA never asks for. We keep those two categories separate internally, and we won't call ourselves “compliant” when “aligned” is the accurate word.
The most useful thing an analyst can say is “I don't know yet.”
Quorum is built to earn the other answers. If that's the kind of health intelligence you've been looking for, we'd like to hear from you.