What would a bad batch look like from where you sit?
The experts were qualified. It passed expert review. QA identified no deficiencies.
Everything you measure can look right while the data is wrong.
For those training the models
The experts were qualified. It passed expert review. QA identified no deficiencies.
Everything you measure can look right while the data is wrong.
The lawyer who wrote your data worked against a timer under conditions they would never accept for client work.
No client would accept them either.
Your model is trained on the result.
The best lawyer you know still gets things wrong.
Credentials improve the odds. They do not validate the data.
More than a bad one.
Not enough to know the data is right.
Reinforcement does not correct it. It does what it was designed to do: reinforces it.
The better the model gets at learning from your data, the more important it becomes that the data is right.
They become part of the data from which the model learns what deserves reinforcement. An error can therefore distort the signal learned from examples that were themselves correct.
A 5% error rate can corrupt far more than 5% of the signal.
One asks whether the output passed inspection. The other asks whether the process was capable of producing the right output.
Training data needs both.
Both are often called QA.
They are not the same thing.
The product is training data that is right.
If your vendor does not guarantee that, it is selling you something else.
Professional legal work carries considerably more. A lawyer can lose the client, the file, their reputation or their licence.
That accountability is valuable, but it is still only an incentive to get the work right.
Decca goes further. We do the work to guarantee that it is right.