Applied research
Will this approach actually hold?
“Which model should we be using?”
A comparison run on your task and your data, with the limits stated, the failure modes named, and a recommendation that says when it stops being true.
A benchmark tells you how something performed on somebody else's corpus. That is a conversation piece and it cannot carry a decision. What can is the same comparison run on your task, with your data, against the boring baseline that quite often wins, reported with the failure modes counted rather than averaged away. An average hides the nine per cent of cases that will generate every complaint you receive. Most of the list below is building a set honest enough that the number means something, then saying when it stops meaning it.
- 01Turn the question into something testable on your task and your data, rather than on a public benchmark.
- 02Build the comparison honestly, including the boring baseline that quite often wins.
- 03Report the limits and the failure modes, not only the headline number.
- 04Say when the finding stops being true: models change, and a result with no expiry is a liability.
- 05Recommend, and be explicit about what evidence would change the recommendation.
- Technical investigationThe question turned into something testable on your data, with the threshold written down first.
- Comparative evaluationEvery approach including the baseline, scored the same way, failure modes named and counted.
- Implementation recommendationWhat to do, what would change our mind, and the date the finding stops being safe to lean on.
A result with no expiry date is a liability.
Make it testable
The question, the threshold, and where the threshold actually comes from.
Build the set
Labelled by hand, awkward cases included rather than excluded, then frozen.
Run and compare
Every approach including the baseline, with failure modes named and counted.
Report
A recommendation, its limits, its expiry, and the evaluation set handed to you.
- Your data, not a public benchmarkA result measured on somebody else's corpus is a conversation piece. It cannot carry a decision, and boards have started noticing.
- A threshold from the processWhat accuracy is genuinely good enough, and who says so. Without it “better” is unfalsifiable and every result is arguable.
- Time from someone who can labelA few hundred examples judged by someone who knows the domain. It is the highest-value input to the exercise and has no substitute.
- Permission to report a negativeSometimes no approach clears the bar. That has to be an acceptable answer agreed in advance, or the exercise is theatre.
What actually is applied research?
A bounded investigation into a technical question you cannot answer from the outside: whether an approach clears your bar, which of several options is actually better on your task, what a method costs at your volume. It runs on your data, with a baseline included and a threshold agreed in advance, and it ends in a recommendation with an expiry date attached.
How is this different from a proof of concept?
A proof of concept shows something can work once. This measures how well, on your data, against the alternatives, with the failure modes named. A demo is not evidence, and boards have started noticing.
What if the result is negative?
You get it, with the evidence behind it. Establishing that an approach does not clear the bar is an inexpensive way to avoid a build that was never going to.
Can you compare vendors for us?
Yes, on your task, measuring the same thing for each. That is usually the reason the vendor claims could not be compared in the first place.
Do we keep the evaluation set?
Always. It is the most durable thing the engagement produces, and it is what lets you re-run the comparison yourself when a new model appears.
How long is a finding good for?
We put an expiry on it. Six months is common for anything resting on a frontier model; considerably longer for findings about the shape of your own data.
If the open question is which problem to solve rather than whether a method works, the prioritisation comes first. Otherwise you get a rigorous answer to a question that did not matter. That is use-case strategy.
Will this approach actually hold?
Describe the question in your own words. We will tell you if it is answerable.