Activity
Mon
Wed
Fri
Sun
Sep
Oct
Nov
Dec
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
What is this?
Less
More

Owned by Samuel

AI Prompting Academy

2 members • Free

Courses & Community where you learn How to Prompt like Boss

Ultimate Content Strategy

161 members • Free

Data Innovators Exchange

757 members • Free

Synthesizer: Free Skool Growth

46.2k members • Free

Data Alchemy

37.5k members • Free

Amplify Views

28.5k members • Free

AI Automation Agency Hub

332.3k members • Free

Content Academy

14k members • Free

IRiS Knowledge Hub

60 members • Free

105 contributions to Data Innovators Exchange
Stop blaming the model for RAG failures 🛑
Most RAG systems are capped long before the model becomes the bottleneck. When a system quietly misses the one crucial passage in production, everyone tends to blame the model—but that gap between "nearly right" and "genuinely trustworthy" is almost never the model's fault. In Part 1 of the new datapro.news series, we are breaking down why your AI's ceiling is actually set before a single prompt even runs. A language model can only reason over what retrieval hands it, meaning no amount of prompt cleverness can conjure back information that was lost three steps upstream. Here are 4 open-source tools that we think change the game: - 🕸️ Crawl4AI: Learn how to bypass web friction and turn messy, JavaScript-heavy sites directly into clean, LLM-ready Markdown. - 📄 Marker: Discover how to process PDFs and documents without destroying their layout, keeping crucial financial tables and LaTeX equations perfectly intact. - ✂️ Chonkie: See how to use lightning-fast semantic chunking so you store whole thoughts, rather than slicing vital sentences or definitions clean in half. - 🗄️ Qdrant: Understand how to combine dense semantic vectors with exact lexical matching (hybrid search) so your database can accurately find specific SKUs, identifiers, and contract clauses. Before you spend another week tuning your model, you need to look hard at what you are feeding it. Check out the video edition below 👇
1
0
Stop blaming the model for RAG failures 🛑
Why your data career feels different in 2026
If your role has felt quietly precarious these last three years, this essay explains why, and it earns the ten minute read. The claim: AI did not eat data engineering, it split it in two. It automated the mechanic and raised the price of the architect, upending two generations at once, the ones who learned the tools and the ones who learned the "why". Both now have to become something new. It is a longer, slower read, built to be sat with rather than skimmed. → check out this weeks edition of www.datapro.news Video edition ⬇️ Which side of it do you see in yourself, and what are you doing about it?
0
0
Why your data career feels different in 2026
Local shell or cloud sandbox: where should your coding agent live?
This week's newsletter digs into the real fork for anyone letting an AI agent near their pipelines, and it is not which model tops the benchmark. It is where the agent runs. Claude Code sits in your shell and harvests dbt metadata locally. OpenAI's Codex runs in a cloud container and needs your data uploaded before it can even start. Both are strong, both come with trade-offs. Which side are you on, and why? Have you let an agent run a dbt build unattended yet, or does that still make you nervous? Full breakdown in this week's newsletter → www.datapro.news · Video hot-take ⬇️
0
0
Local shell or cloud sandbox: where should your coding agent live?
The Three-Tier Stack Just Got a Lot Harder to Defend.
If you're still designing your stack with a separate transactional layer, an analytical warehouse, and a real-time tier stitched together by pipelines — H1 2026 just made that architecture harder to justify. Every major platform shipped something this half that chips away at the need for that middle layer. This week's issue breaks down exactly what ➡️ Snowflake, ➡️ Databricks, ➡️ BigQuery, ➡️ Redshift, and ➡️ MS Fabric each did — and what the pattern means for the rest of the year. Video version of this weeks www.datapro.news below 👇
The Three-Tier Stack Just Got a Lot Harder to Defend.
The job description for data engineers quietly changed this month.
For a decade our job was to make capability flow: Get the data in, get the model serving, ship the pipeline. After Washington switched off Claude Fable 5 for half the world overnight, there's a second mandate now - make capability survivable. ➡️ Designing for the provider disappearing. ➡️ For the model silently getting worse with no error. ➡️ For the legal status of your output being genuinely unsettled. Multi-model routing and dependency governance just moved from "nice architecture" to core competency. The engineers who can answer "what happens to this pipeline when the model goes dark?" are going to be the ones writing the architecture decisions for everyone else. Check out the video edition below 👇
0
0
The job description for data engineers quietly changed this month.
1-10 of 105
Samuel Williams
5
222 points to level up
@samuel-williams-6637
Checking this out

Active 5d ago
Joined Apr 8, 2024
Powered by