Activity
Mon
Wed
Fri
Sun
Nov
Dec
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
What is this?
Less
More
3 contributions to Brendan's AI Community
800,000 Lines of Unit Tests, Deleted. Smart or Reckless?
A startup just deleted more than 800,000 lines of unit tests. On purpose. The founder of an AI observability startup argues that now that coding agents write so much of the code, unit tests can do more harm than good: → They lock in bad AI-generated code → They create "busy work" for agents → They slow down CI pipelines → They drive up token bills A Coinbase engineer agreed: lean on full end-to-end, integration, and golden tests instead. Not everyone is convinced. Veteran engineer Bob Martin ("Uncle Bob") says he still relies on unit tests to keep his agents honest. Here's how I see it: Unit tests written by agents, for code written by agents, can turn into the AI grading its own homework. But tests that capture real business rules are the guardrails that stop an agent from "fixing" your code into the wrong shape. So maybe the question isn't "delete or keep?" It's "which tests protect behavior, and which only protect implementation details?" Where do you stand? A) Delete most of them and test behavior end-to-end B) Keep them, because they're the agent's guardrails
800,000 Lines of Unit Tests, Deleted. Smart or Reckless?
1 like • 9h
Unit tests related to guardrails should never be deleted. Tests generated by an AI agent are akin to the agent writing its own code and then reviewing it with its own unit tests; therefore, such tests should be considered for deletion after review.
AI Harness Tax
Same model. Same task. Same success rate. Half the bill. A new study from UC Berkeley and Arena put a name to something many of us building with AI agents have felt but rarely measured: the harness tax. The harness is the agent wrapper around the model: Claude Code, Codex CLI, Pi, and others. It turns out the harness can change your inference costs as much as the model you pick. A few findings stood out: → Claude Fable 5 solved ~97% of tasks in Claude Code, Codex CLI, and Pi. But Claude Code cost $1.33 per rollout, while Pi cost $0.67. → On SWE-bench Lite, Claude Code cost about 2x more than Pi and 1.6x more than Codex. Success rates were within 2 percentage points of each other. → Much of the gap appears before the agent does anything. Claude Code starts with 27,000+ tokens of context, compared with ~2,000 for Pi. → Models don't always perform best inside their own vendor's harness. Simple setups are often surprisingly competitive. The most interesting part for me: even when two setups post identical benchmark numbers, they often fail on completely different tasks. A leaderboard score won't tell you which one fits your codebase. The researchers' advice is practical: 1- Test a few model and harness combinations on your actual engineering workload 2- Measure cost per solved task, not raw cost per run 3- Pick the cheapest option that clears your reliability bar 4- Retest whenever the model or harness changes If you're running agents at scale, whether for your own product or for clients, this is the difference between a margin and a leak. The best agent setup isn't the most elaborate one. It's the one that solves your problems at a cost you can sustain.
AI Harness Tax
💰 $5000 Voice AI Hackathon! (Feb 25th)
We're launching a $5000 Voice AI Hackathon, in partnership with the amazing team at Retell AI! This is a 7-day sprint to build voice agents that solve real business problems. What’s the point? → Build and demo a working AI voice agent using Retell's cutting-edge tech. → Compete for a $5,000 prize pool → Get hands-on experience building the next generation of voice AI. → Showcase your build to a community of agency owners, founders, and automators. This is for AI agency owners, automation builders, SaaS founders, and anyone excited about the future of voice AI. 🥇 1st Place: $1,000 cash + $1,000 in Retell credits 🥈 2nd Place: $500 cash + $750 in Retell credits 🥉 3rd Place: $250 cash + $500 in Retell credits 🏅 2 Honorable Mentions: $500 Retell credits each Want extra build support? Launchpad members get access to our dev team, reviews, and troubleshooting during the hackathon. The official kickoff is on February 25th @ 2PM EST. Launchpad members receive a 48-hour early build window starting February 23rd, with access to setup resources and dev support. To participate: Comment 'HACKATHON' I will send you the registration link via a direct message here.
💰 $5000 Voice AI Hackathon! (Feb 25th)
0 likes • Feb 14
HACKATHON
1-3 of 3
Waleed Ijaz
3
4 points to level up
@waleed-ijaz-6248
Building AI Agents and Automations for Businesses to increase revenue and reduce time

Active 5h ago
Joined Jun 30, 2025
Powered by