Nephila Radar [2/26 | Sep 14 – Sep 27]
On our radar
Addy Osmani, an engineer and Member of Technical Staff at Anthropic where he works on Claude Code, puts a name to the problem hiding behind codebases that look clean: comprehension debt, the growing gap between how much code exists in a system and how much of it anyone actually understands. He cites an Anthropic study of 52 engineers: those who used AI to generate code finished the tasks in about the same time as those who didn't, but scored 17% lower on a follow-up comprehension quiz (50% versus 67%), with the sharpest drop in debugging.
Worth reading to understand why reviewing AI-generated code takes a different discipline than reviewing code written by people: Comprehension Debt - the hidden cost of AI generated code - Addy Osmani
A LeadDev panel with Birgitta Böckeler, Global Lead for AI-assisted Software Delivery at ThoughtWorks, and Grace Torany, DevOps Engineer at NVIDIA, looks at the flip side of AI-assisted development. Research cited in the panel also shows that engineers who adopt AI aggressively, without full understanding or verification, end up with a sense of speed they don't actually have.
Worth watching to figure out where to set the guardrails before the technical debt created by agents turns into a security or reliability problem: How to clean up the tech debt your AI is creating - Grace Torany, Aaron Newcomb, Birgitta Böckeler.
Requires free registration (name, email) on the LeadDev platform.
Ajeet Singh Raina, Developer Advocate at Docker, tells the story of an incident from December 2025: a developer asked Claude Code to clean up an old repository, the agent generated the command rm -rf tests/ patches/ plan/ ~/, and the trailing slash after the tilde wiped the entire Mac home directory, with no way to recover it because SSD TRIM had already zeroed the freed blocks. It wasn't an attack: the agent did exactly what it was asked to do, but no architectural boundary stood between the generated command and its execution.
Worth reading to understand why running an agent with a user's full permissions, on the same filesystem as their personal data, is a structural risk and not an isolated bug in one tool: Coding Agent Horror Stories: The rm -rf ~/ Incident - Ajeet Singh Raina
From the Nephila blog
At PyCon Italia 2026, Ines Panker observed in her talk that asking for an estimate is often really about reducing the risk of a decision, not just finding out how long something will take. Starting from that idea, Emanuela Dal Mas, Managing Director at Nephila, explains how we fit estimates into a broader information stack made of assumptions, confidence levels, and historical data, with two concrete practices: peer double-review on stable teams, and planning poker when skills are more mixed.
Worth reading to see how we try to make estimates more reliable, so we can steer the future one iteration at a time: Beyond estimates: building an information stack for risk management - part 1 - Emanuela Dal Mas
Try it: find the technical debt hiding in your codebase
Fallow reads a TypeScript or JavaScript project's entire module graph with a single command, npx fallow, no configuration needed on the first run, and flags dead code, duplication, complexity hotspots, and drift from the architecture or design system. It also integrates with agents like Claude Code and Cursor via MCP, so the agent can query it directly while it works, instead of just suggesting a change in chat.
Worth trying on a TypeScript or JavaScript project where you suspect technical debt has piled up: Fallow
***
Did you find this list of resources interesting, or want to share your thoughts? Join the conversation on our LinkedIn page, sharing ideas and comparing perspectives is the part we enjoy most.
***
We wrote this article with the support of AI. If you’re interested in how we use these tools as part of our writing process, you can read more here.
