<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Coding-Agents on Kuldeep Pisda</title><link>https://kdpisda.in/tag/coding-agents/</link><description>Recent content in Coding-Agents on Kuldeep Pisda</description><generator>Hugo</generator><language>en-us</language><lastBuildDate>Mon, 14 Sep 2026 09:00:00 +0530</lastBuildDate><atom:link href="https://kdpisda.in/tag/coding-agents/index.xml" rel="self" type="application/rss+xml"/><item><title>Real-SWE: The Benchmark Where Coding Agents Meet an Actual Enterprise Codebase</title><link>https://kdpisda.in/real-swe-benchmark-frontier-coding-agents-enterprise-codebases/</link><pubDate>Mon, 14 Sep 2026 09:00:00 +0530</pubDate><guid>https://kdpisda.in/real-swe-benchmark-frontier-coding-agents-enterprise-codebases/</guid><description>&lt;p&gt;Every few months a new coding benchmark claims a frontier model &amp;ldquo;writes production code.&amp;rdquo; Then someone runs that same model against an actual production codebase, one with tax rules, half-documented billing logic, and three ways of doing the same thing because three different engineers touched it over five years, and the score falls off a cliff. That&amp;rsquo;s exactly what happened this week. Specific Labs, a YC-backed startup, published Real-SWE, a benchmark built entirely from private codebases it licensed from real companies, including a consumer events app with 200K+ users and a fintech platform processing 100K+ bank statements. The best model on it resolves 38.8% of tasks. On SWE-bench Verified, frontier models routinely clear 70-90%. That gap is the story, and it matters to anyone deciding whether to let an agent loose on their own repository rather than a curated GitHub issue.&lt;/p&gt;</description></item></channel></rss>