TilliT
Check a product
Human Review Process

How our human review follows the scientific method

A badge on a product page is, in effect, a claim: "we checked this, and here's what we found." When we designed TilliT's human review process, we borrowed on purpose from fields that have spent decades answering exactly that question: clinical trial methodology, systematic review, psychometrics and inter-rater reliability research, and research ethics. Here are ten specific ways that shows up. Not as a metaphor, as load-bearing process.

1

A falsifiable prediction, before review begins

The scientific method: a hypothesis stated precisely enough to be proven wrong.

Before any human reviewer looks at a case, our system produces a preliminary status from a fixed, published formula. That's a real, specific, checkable prediction, recorded before human review, so we can later ask whether the reviewer agreed with it or overturned it, and why.

2

Blinded, independent replication

The scientific method: an independent lab repeating an experiment without seeing the original result.

A random sample of decided cases is routed to a second reviewer through a genuinely blind process. They see the same underlying evidence, but never the first reviewer's identity, verdict or reasoning.

3

Formal inter-rater reliability measurement

Psychometrics: agreement between raters is measured, not assumed.

From every blinded second review we compute percent agreement and Cohen's kappa. It's the same statistic used in clinical and behavioral research, chosen because it corrects for the agreement you'd expect by chance alone.

4

A second opinion on the highest-stakes findings

Peer review: a significant result is confirmed by someone other than the person who produced it.

A correction that swings a score dramatically, or moves a finding into or out of a serious conflict, applies immediately but carries a visible “pending confirmation” state until a second, genuinely different reviewer has independently checked it.

5

Disclosed conflicts of interest

Research ethics: a reviewer with a stake in the outcome discloses it, or recuses.

Every reviewer declares any financial, employment or personal relationship with a brand before they're allowed to review that brand's case. It's enforced automatically, and declarations get refreshed on a standing schedule.

6

A versioned, published protocol

Clinical trial methodology: a trial follows a registered protocol, and any amendment is a tracked, versioned event.

Every assessment carries an explicit methodology-version stamp. If the underlying rules change, the version changes with them, and an automated check refuses to let the two drift apart silently.

7

Reproducibility, backed by a regression library

The scientific method: a result should be reproducible by re-running the experiment.

The same evidence always produces the same score. Before any change ships, it has to pass 45 real, worked test cases spanning every badge, so a fix for one situation can't silently change the answer for another.

8

Continuous self-correction, not a one-time judgment

The scientific method: a theory is provisional, and stands only until better evidence overturns it.

An automated process periodically reviews rules that keep getting overturned by reviewers, on the premise that a rule people keep disagreeing with is more likely to be the rule that's wrong. Corrections are always layered on top of the original record, never a silent rewrite.

9

Watching the instrument, not just the results

Experimental rigor: a fatigued or overloaded observer is a known source of measurement error.

Reviewer workload is tracked explicitly. A day of unusually high volume for any one reviewer gets flagged for attention. It's the same reasoning behind duty-hour limits for pilots and shift-length limits in hospitals.

10

A complete, permanent chain of custody

Systematic review & financial audit: every step is documented well enough for someone else to retrace it.

Every fact gathered, every reviewer decision, and every correction's stated reason is recorded in an append-only log. Nothing is ever edited or deleted. You can only add to it.

Where the analogy has limits

We think it would be dishonest to lean on "scientific method" language without being precise about where the comparison stops.

Operational rigor, not a peer-reviewed publication

Everything described here is real, internal practice: measured, enforced, logged. It's not the same as a study published in and vetted by an external scientific venue, and we don't present it as one.

Blinding is procedural, not absolute

Our double-review process withholds the first reviewer's identity and verdict by design. But, like any blinded process run by a small team, it depends on reviewers not talking about an active case informally. We treat that as a real, current limitation, not a solved problem.

Falsifiable predictions, blinded replication, measured agreement, disclosed conflicts, and a permanent record are how science makes a judgment call trustworthy to someone who wasn't in the room. We built our review process around those answers on purpose.