Every evaluation-driven development pitch lands the same way: write the eval first, the way TDD writes the test first, and let the score gate the release. The suite genuinely is better than the demo and the product manager's intuition it replaced. The trouble starts about two weeks after adoption, when someone asks whether a passing build means anything a user would recognise as quality — and the answer turns out to be that nobody knows, because the judge producing the number was never scored against a human. In this edition of datapro.news, we cover the step the announcements skipped. EDD has a real academic anchor: CSIRO's Data61, with UNSW and ANU, published the reference treatment in November 2024 and revised it as recently as November 2025, and its EDDOps model treats evaluation as continuous governance rather than a terminal checkpoint. Read the definition carefully, though. Evaluation there is an umbrella covering testing, benchmarking, validation, monitoring and risk assessment — it is a process model, not a metric. It tells you where evaluation belongs, not what good is. The load-bearing component in "the evals passed" is a judge nobody has graded. Here are the 3 layers the decision actually comes down to: 🔬 Error analysis: Learn why the industrial methodology inverts the order most teams use — read real traces, open-code the failures by hand, cluster into a taxonomy, count frequencies, and only then build evaluators. Plus the UIST 2024 finding on criteria drift that quietly kills the clean TDD analogy, since grading outputs is how people discover their criteria, and the two contrarian positions held by the practitioners with the most reps. ⚖️ LLM-as-a-judge: Discover the fair version of the problem rather than either caricature — over 80% agreement with humans in the foundational MT-Bench work, matching human-human levels, alongside documented position, verbosity and self-enhancement bias. The most careful position-bias study ran 15 judges across 22 tasks and 150,000-plus evaluation instances, and found the bias worsens as the quality gap between candidates narrows, which is exactly where your regression tests live.