Claude's fingerprints now mark ~40% of 'human' GitHub pull requests
Louis Abraham’s Show HN is an interactive chart built on a daily scrape of about 1,000 GitHub pull requests. Across roughly 461,000 PRs and 51 million words collected from January 2025 through mid-August 2026, it sorts PR vocabulary into ten clusters with KL-divergence k-means and tracks each cluster’s weekly share. One cluster is the whole story: it barely existed in early 2025 (about 0.7% of PRs) and by last month accounted for close to 40% of pull requests attributed to human authors.
That cluster’s most representative words are the verbal tics of Claude and other coding agents — terms like load-bearing, plainly, quietly, refusal, re-derived, genuinely, deliberately, byte-identical, and carries. In other words, you can now detect agent-written contributions from word choice alone, and their share of nominally human open-source work is both large and climbing fast. The remaining clusters map other recognizable dialects: automated bots (Renovate, Snyk, qodo), CI and package-manager flag jargon (—locked, —frozen, —test-threads), non-English descriptions, and an older 2025-era hedging register (seems, basically, probably, unfortunately) that is now shrinking as the agent cluster expands.
The significance is provenance. A rapidly growing fraction of code changes credited to people is really being drafted by LLMs, and a simple corpus-linguistics method makes that visible at scale. For anyone weighing the trust and authenticity of open-source contributions — who actually wrote the code getting merged — this is a concrete, quantified snapshot of how quickly agent vocabulary has taken over the pull-request stream.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.