Activity
Mon
Wed
Fri
Sun
Nov
Dec
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
What is this?
Less
More

Owned by Lina

Exclusive group for DataVault4dbt Premium members: access tutorials, premium content, community Q&A with experts, and a calendar for live sessions.

40 contributions to Data Innovators Exchange
Your AI doesn't know what "active customer" means. It will answer anyway. 🎯
Ask three systems how many active customers you have and you get three answers: 5,892 from the Sales CRM, 6,155 from the Marketing DB, 6,001 from the Finance Ledger. Sales counts recent logins. Finance counts accounts with an open balance. The old warehouse relies on status flags nobody has updated in years. Then you put an AI agent on top, and it blends the rules or quietly picks one at random. The failure doesn't look like a failure: a clean chart, a plausible number, a green "task completed," and no warning that the number is wrong. In this edition of datapro.news, we look at why the gap between AI demos and AI you can trust is a semantic problem, not a model problem. Gartner's numbers set the stakes: 60% of AI projects without AI-ready data are expected to be abandoned, and 63% of organisations either lack the right data management practices for AI or aren't sure they have them. The model can be as capable as you like. If business meaning was never written down, it is guessing. Here are the 3 layers the fix actually comes down to: 🧭 The decoder ring. Learn the difference between the three layers of meaning that most teams blur together: a taxonomy gives you the vocabulary (what a SKU is), an ontology gives you the concepts and how they link (what an Order is and what it connects to), and a semantic layer gives you the rules (what "Active" means, and who decided). The evidence is striking. In published benchmark work, GPT-4's accuracy on enterprise questions rose from 16% against a raw schema to 54% when given an ontology. It's the same model with the same data, and the only change is that the meaning was made explicit. The semantic layer is also where enterprises most often stall, because that is where conflicting source definitions finally have to be reconciled. 🏛️ The Data Vault backbone. See why Data Vault is a natural home for business meaning, because its structure echoes an ontology. Hubs hold the core concepts' business keys, Links record the relationships, and Satellites carry the descriptive detail. The Raw Vault stores data exactly as sources deliver it, so every disagreement stays traceable. The Business Vault then resolves the conflict with one agreed rule, instead of leaving it to whichever dashboard or agent happens to run the query.
0
0
Your AI doesn't know what "active customer" means. It will answer anyway. 🎯
The measurement war is over. Nobody was assigned the part that checks the measurer. 📊
Every evaluation-driven development pitch lands the same way: write the eval first, the way TDD writes the test first, and let the score gate the release. The suite genuinely is better than the demo and the product manager's intuition it replaced. The trouble starts about two weeks after adoption, when someone asks whether a passing build means anything a user would recognise as quality — and the answer turns out to be that nobody knows, because the judge producing the number was never scored against a human. In this edition of datapro.news, we cover the step the announcements skipped. EDD has a real academic anchor: CSIRO's Data61, with UNSW and ANU, published the reference treatment in November 2024 and revised it as recently as November 2025, and its EDDOps model treats evaluation as continuous governance rather than a terminal checkpoint. Read the definition carefully, though. Evaluation there is an umbrella covering testing, benchmarking, validation, monitoring and risk assessment — it is a process model, not a metric. It tells you where evaluation belongs, not what good is. The load-bearing component in "the evals passed" is a judge nobody has graded. Here are the 3 layers the decision actually comes down to: 🔬 Error analysis: Learn why the industrial methodology inverts the order most teams use — read real traces, open-code the failures by hand, cluster into a taxonomy, count frequencies, and only then build evaluators. Plus the UIST 2024 finding on criteria drift that quietly kills the clean TDD analogy, since grading outputs is how people discover their criteria, and the two contrarian positions held by the practitioners with the most reps. ⚖️ LLM-as-a-judge: Discover the fair version of the problem rather than either caricature — over 80% agreement with humans in the foundational MT-Bench work, matching human-human levels, alongside documented position, verbosity and self-enhancement bias. The most careful position-bias study ran 15 judges across 22 tasks and 150,000-plus evaluation instances, and found the bias worsens as the quality gap between candidates narrows, which is exactly where your regression tests live.
0
0
The measurement war is over. Nobody was assigned the part that checks the measurer. 📊
The specification war is over. Nobody was assigned the part that runs. 📄
Every data contract pitch lands the same way: one YAML file, agreed between producer and consumer, versioned in Git, executable. The file genuinely is better than the wiki page it replaced. The trouble starts about two weeks after adoption, when someone asks what happened when the freshness SLA was missed on Sunday — and the answer turns out to be nothing, because nothing was watching. In this edition of datapro.news, we cover the question the announcements skipped. ODCS v3.1.0 shipped in December 2025, the competing Data Contract Specification was deprecated by its own maintainers, and Bitol graduated at LF AI & Data in July 2026. Authoring in ODCS is now the boring, correct default, which is the highest compliment a standard can receive. But ODCS is a declarative specification, not an execution engine — it defines the what and deliberately leaves the how and the when to whatever system processes the data. Quality rules describe what should be true; something else has to go and check. The word doing the most work in "executable data contracts" is not contract. Here are the 3 layers the decision actually comes down to: - 📐 ODCS v3.1.0: Learn what actually shipped — relationships that hold even where the store enforces nothing, SLAs that carry a schedule, a registered media type — and the asterisk on the backward-compatibility claim, because "no migration required" and "run the linter before you upgrade CI" are different instructions and only one of them is accurate. - 🏛️ Bitol governance: Discover why graduation is the fact that should drive a procurement decision rather than any feature list, since it is a checklist and not a sentiment — plus the two things to verify yourself, including a foundation page that still says incubation and adoption figures published by the standard's own chair. - ⚙️ The enforcement layer: See why choosing ODCS is now low-risk and choosing what runs it is not. Soda Core changed licence in January 2026, GX Core stewardship passed to Fivetran in May, and the dbt Labs merger closed on 1 June. The open spec consolidated at exactly the moment the engines beneath it did.
The specification war is over. Nobody was assigned the part that runs. 📄
The multimodal lakehouse is real. The diagram everyone is drawing is not. 🎞️
Every RAG project starts the same way: a few hundred PDFs, a chunking strategy copied from a blog post, an off-the-shelf vector store. The demo genuinely works and the budget clears. The trouble starts a quarter later, when the business stops asking about documents and wants the model to review the security footage. In this edition of datapro.news, we cover the rewrite nobody scoped. Unwinding a text-only stack is a real architectural project with real money behind it — but the reference diagram circulating to describe it (lakehouse, Kafka, Flink, Materialize, done) does not survive contact with the documentation of its own components. Three of the four boxes are doing jobs they do not do. Here are the 3 components the design actually turns on: * 📼 Apache Kafka: Learn why "Kafka ingests the video stream" tells a design review you have not sized a broker, and the pattern that already has a name you should be using instead. * ⚡ Apache Flink: Discover why the part of the pitch that sounds most like vapour became true most recently — plus the two footnotes the marketing omits, one of which means you will pay for some tokens twice. * 🧮 Materialize: See why listing it as a way to compute embeddings is the tell that a diagram was assembled from vendor landing pages, and where it genuinely belongs instead. We also get specific about the distinction vendors blur between Lance and LanceDB, why "batch is deprecated" is contradicted by the flagship deployment of this very architecture, and the compliance citation that will get your design second-guessed by legal. The issue closes with a scorecard: what each component does, and what to check before you commit. Parsing a PDF is a commodity, and it was one before this wave. The scarce thing is a system that embeds a live stream without a bespoke microservice holding it together and joins the result against governed relational data. That is buildable in 2026 — just not from the diagram everyone is drawing. Get the boxes right first.
0
0
The multimodal lakehouse is real. The diagram everyone is drawing is not. 🎞️
The format war is over. Your next lock-in moved upstairs. 🧊
Every Iceberg pitch lands the same way: your data, your object storage, an open format, queryable from anywhere. The demo genuinely works. The trouble starts about two weeks into production, when someone asks who is allowed to see column seven — and the answer turns out to live somewhere that isn't open at all. In this edition of datapro.news, we cover the fight nobody announced. A directory of Parquet files in a bucket is inert; it is storage, not a database. Something has to tell an engine which metadata pointer is the current valid state of a table, and whether the querying principal may read it. That something is the catalog. By commoditising the file format, the industry did not eliminate the vendor lock-in it spent a decade complaining about — it relocated it, from the storage layer where it was visible and much-discussed to the governance layer where it is neither. Your Parquet files stay portable. The 5,000 grants, masking rules and policy tags you authored do not. Here are the 3 catalogs the decision actually comes down to: - 🧭 Apache Polaris: Learn why graduating to Apache Top-Level Project in March 2026 is worth more to you than any feature comparison, and why the neutrality you get with it is also an on-call rota — self-hosting means owning HA for the one service every pipeline and dashboard in the company stops without. - 🌿 Project Nessie: Discover the most elegant idea in the field — Git branching, tagging and merging across your entire lakehouse state — plus the due-diligence detail that should stop you adopting it as a destination, because its own sponsor has stated it will fold those capabilities into Polaris and retire the project. - 🔗 Unity Catalog: See why the criticism you have probably repeated in an architecture review is now out of date, and where the real asymmetry sits instead: outbound it is a well-behaved Iceberg REST server, inbound a reluctant client, and that is a design choice rather than a missing feature.
The format war is over. Your next lock-in moved upstairs. 🧊
1-10 of 40
Lina Sibbel
4
5 points to level up
@lina-sibbel-7665
Marketing Associate @Scalefree

Active 2d ago
Joined Apr 11, 2024
Hanover, Germany
Powered by