When a sensor hands you a number, the number is the easy part. The hard question is how far to believe it.
Most of us answer with adjectives: “accurate”, “robust”, “validated”, “medical-grade”. They describe how a device behaved on the day someone tested it. They say nothing about the day it is strapped to a patient who sweats, or left in a field through a wet season.
Readiness scales stop us fooling ourselves about progress. A technology readiness level asks whether something functions at a given stage. It does not ask whether you can trust it when the world moves, and in sensing that gap does the most damage.
So I developed a scale for trust itself, the Validation and Trust Readiness Level, or VTRL™. It has six levels across three stages: bench, field and replication. I define each separately, because pairing them hides the distinction that matters most.
Level 1: characterised. On the bench, you identify a reference comparator, something you already believe, against which the sensor can be judged. You build an uncertainty model and quantify it.
Level 2: failure-aware on the bench. The sensor passes a drift test under controlled conditions, and is shown to detect its own calibration failing.
Level 3: observed in the field. The sensor leaves the bench and the comparator goes with it, sitting beside the sensor in the real setting. You record how the device fails there: where it drifts, which conditions break it, how it disagrees with the reference. At this level the sensor is studied, not relied on.
Level 4: tested against its own warnings. Now the sensor’s signal-quality score is judged against thresholds fixed in advance, on failures and conditions it was not tuned on. Does it flag the readings that were wrong? Does it stay quiet on the ones that were right? Missed warnings and false alarms both count against it. Ethical clearance is in place where people are involved.
Level 5: replicated. A second site, a different operator, and the variation between the two measured.
Level 6: accountable. That variation stays within stated bounds, the data are governed and archived, and the result is published.
Here is why the gate sits at level 4. At level 3 you learn how the device fails. At level 4 you learn whether it recognises failures it has not seen. Only then is deployment defensible, and only within the conditions tested, with ongoing spot checks against a reference. Nothing is called deployment-ready below level 4. Until then, any claim of readiness is a hope.
The reference comparator is the heart of this, and it hides a difficulty. A sensor deployed alone has no reference to check itself against. Telling that a measurement is wrong without something to compare it with is the hard research problem.
The scale makes it tractable. At level 3 you place the reference beside the sensor so that you can learn how the sensor looks when it is going wrong. At level 4 you test whether its own warnings catch failures it was not tuned on. At level 5 you ask whether those warning signs survive a change of place and operator. The goal is a sensor that has learned its failure signatures, so that it can say “do not trust me right now” when no reference is present.
Four principles apply at every level. They are my proposals, not an established procedure.
First, confidence should travel with every reading. A number without a statement of how far to believe it should be treated as incomplete.
Second, false alarms count as failures at every level. A fall detector that cries wolf until a caregiver switches it off has failed as surely as one that stays silent during a fall. A trust level should include how often the warnings are wrong.
Third, every claim should come with an operating envelope: the temperature, humidity, contact and use conditions within which the level was earned. Outside the envelope, the level does not apply.
Fourth, a level should be revocable. Change the material, the firmware or the site, or let drift pass its stated bounds, and the device drops back until it re-earns the level. Trust that cannot be lost is a label.
Two of my demonstrations show the early levels. A 2025 paper I co-authored describes a device that monitors a drug in freely moving animals and applies baseline correction and drift compensation on the device itself. In a 2025 plant-sensor paper, a cyclical cleaning step extended the sensor’s lifespan by preventing electrode passivation, and its readings were compared with mass spectrometry. By this scale, these two demonstrations sit mostly at levels 1 to 3, and neither paper claims more. That describes those demonstrations, not the rest of my research.
This will make some claims smaller. A device that looked finished will turn out to be at level 2. Papers will have to say so, and funders will have to accept it. The people who depend on these devices, patients, farmers, building managers, are already carrying the cost of overstated claims. They never see the bill.
I do not want sensors that are never wrong. No sensor is. I want sensors that have been caught being wrong, in the open, under the conditions that matter, and that know how to tell us. Trust earned that way is a level you can point to, and it is one you can lose.
Papers mentioned
- A portable device for real-time continuous drug monitoring in freely moving small animals, Biosensors and Bioelectronics, 2025
- In vivo dynamics of indole- and phenol-derived plant hormones: long-term, continuous, and minimally invasive phytohormone sensor, Science Advances, 2025