Distribution Is Outrunning Verification in the Coding-Agent Skills Layer
Published October 9, 2026 · 2041 words · 11 min read
A repository that teaches coding agents to mod Terraria, Age of Empires II and GTA V has no obvious claim on the attention of VCs and CTOs. Yet rehan-remade/universal-modder is a visible public example of a pattern forming around coding agents: the channel that distributes agent capabilities is scaling faster than any way of checking what travels through it. The repository appeared on GitHub on September 30 and, per AI Frontier Post, reached about 1,430 stars within a day and a half of launch. According to GatiFlow intelligence, our GitHub collector first logged it at 2,917 stars on October 4 and last read 5,170 on October 7, a rise of 77.2 percent across four tracked days. New stars were still arriving at a steady daily pace, with no visible decay. On its own that is a single GitHub signal. The data around it is what gives it meaning.
The contents matter more than the count. Universal-modder bundles what an agent needs to research, build and test a mod, according to its Hugging Face guide and AI Frontier Post. The kit has skills in the Agent Skills format, a command-line tool called um, an MCP server for generating art, 3D models and audio through fal, and a shared knowledge base of field notes on how specific games were modded. It installs through the plugin marketplaces of Claude Code and Codex CLI, as a Gemini CLI extension, through VS Code and Copilot, or as a skills-only package for any compatible agent. It was built for the open channel rather than for one vendor.
The loop it enforces is the part worth copying. The agent backs up saves first and asks before driving the mouse or keyboard, installing a loader or publishing anything. It tests in the running game instead of trusting a clean build, then runs a publish check that blocks game files, decompiled code and leaked keys. AI Frontier Post argues that the field-notes knowledge base is the real product and that proving a mod works in the live game is where these projects are won or lost. The rest of the data points the same way.
The channel itself widened on October 1, when Anthropic released Mods in Claude Code v2.1.287. A mod is a plugin whose code runs inside the Claude Code process with the permissions of the person who installed it, and mods load by default from that version, per Anthropic's documentation. Skills package procedures, MCP servers expose tools, and mods change how the host itself behaves, so each carries a different trust burden. Universal-modder's own components are skills, a command-line tool and an MCP server rather than mods, and it ships through the same plugin channel. AgentConn reported 833,000 views for the launch post from Boris Cherny, head of Claude Code, and counted ten of the fifteen repositories on GitHub Trending on October 3 as agent skills, harness optimizers or orchestration tools.
The window that matters is the next two to four weeks, not this week's news. The standard underneath is old enough to have a market and young enough to lack governance. Agent Skills came out of Anthropic in October 2025 and the specification was opened up that December; by June, about 40 products appeared on the standard's showcase, per Agentman's ecosystem report. The directories are now large. The claudemarketplaces.com index advertises more than 23,600 Claude Code skills. Agentman, citing Agensi, counted more than 2,500 plugin marketplaces registered there in June, and it puts the SkillsMP aggregator at about 1.9 million skills scraped from GitHub. Those are index sizes, not vetted inventory. A channel that fills this quickly gets its defaults set before anyone audits them.
The architectural choice looks like open versus curated, but the numbers show a vendor shelf next to an open bazaar. claudemarketplaces.com lists three plugins in Anthropic's public skills repository and 222 in its official claude-plugins-official marketplace. The rest of the indexed catalog comes from third parties, from Vercel and Microsoft to individual developers. The open-standard bet buys portability: a skill written once runs in Claude Code, Codex, Cursor, Gemini CLI and Copilot, so competition moves to quality and distribution. The curated bet buys control. Between them sits a third option, a private marketplace that an organization runs itself.
The control surface already exists for those who look. Anthropic's documentation describes managed settings, including allowManagedModsOnly, that let administrators limit loading to the organization's own mods and the ones built into Claude Code. Its overview also says claude plugin validate previews a mod's event hooks and the actions it requests, including file reads and network calls, without running any of its code. For an engineering leader the decision is therefore less open versus closed than who may put code on a developer machine, and whether that policy exists before the first enthusiastic install.
The registries show the substrate moving. According to GatiFlow intelligence, weekly downloads of @anthropic-ai/sdk on npm reached 54.0 million on October 7, up 12.8 percent from 47.9 million a week earlier, while openai's package rose 10.0 percent, from 46.2 million to 50.8 million. Two rival SDKs growing together suggests more codebases wiring in agents through whichever vendor is at hand, though download counts include automated installs and say little about what runs in production.
GitHub adds two reference points. Louis-CFM/coucou, a notch-based companion app that lets developers approve Claude Code permission prompts, also supports Codex and Cursor and stood at 3,850 stars when our collectors last read it on October 7. It is tooling for supervising agents rather than building them. Our collectors also flagged storytold/photocraft, a Rust reimplementation of Photoshop with no agent angle, which more than doubled its stars within a single day of tracking. Star velocity measures attention before it measures adoption.
The research layer contributed three papers that our collectors surfaced together on October 6, all submitted to arXiv on October 2. One, accepted at NeurIPS 2026, finds that tool-using multimodal models refuse harmful requests less often than the same models without tools, with the relative failure rate rising by up to 68.7 percent. A second, from IBM authors, measures the price of agent-to-agent protocols. On one test machine the A2A SDK spent 62.71 milliseconds per hop on connection setup against 2.56 for NLIP, a gap that narrows to near parity with connection caching on faster hardware. The authors attribute part of the cost to the SDK's implementation rather than the protocol. A third describes a confidence layer that decides at three points whether a clinical text-to-SQL agent should stop: before it runs, mid-run, or before it delivers an answer. None concerns Agent Skills directly. They probe the layers around skills: safety under tool use, the cost of coordination and when an agent should abstain.
The production evidence for the same gap sits inside the skills ecosystem. A January 2026 study analyzed 31,132 skills from two marketplaces and found that 26.1 percent contained at least one vulnerability, with skills that bundle scripts 2.12 times as likely to be flawed as instruction-only ones. SkillsBench, a February 2026 benchmark, found that curated skills raised agent pass rates by 16.2 percentage points on average, that self-generated skills gave no average benefit, and that 16 of its 84 tasks got worse with curated skills. A skill can bundle scripts that run with the user's permissions, and the safety paper is a reminder that adding tools can change model behavior in ways a test without tools would not predict.
AgentConn frames the Mods launch as an app-store moment and names security as the open risk. The measurements point to a second gap, plain effectiveness. SkillsBench found its smallest gain from curated skills in software engineering, 4.5 points, against 51.9 in healthcare, which fits the idea that skills pay most where the base model knows least. Agentman's table of Skills.sh categories shows supply running the other way, with 288,811 development and engineering skills against 6,354 in healthcare and life sciences. Attention and volume sit where the measured uplift is thinnest. What makes universal-modder interesting is therefore not the nuke in Terraria but version-specific field notes and a live game as the arbiter of success. That is niche, curated knowledge with a built-in check, closer to the profile where the benchmark found its largest gains than to generic coding advice. Our read: the market is over-rotating on catalog size and under-pricing the verification harness and the curated knowledge behind a skill.
Hiring data adds a thinner layer. daily.dev Recruiter's June guide says postings for AI agent developers rose 340 percent year over year as of early 2026. AY Automate argues that few engineers have more than two years of hands-on MCP experience because the specification only shipped in late 2024. Both vendors sell into this market. Our collectors give a first-party check on demand: they flagged a senior product manager role for an enterprise MCP and AI control plane at Workato and a forward-deployed engineer role on agent runtimes and MCP at NMK Global. The roles exist. The evidence that talent is scarce remains anecdotal.
If you are building developer tools, an agent product or an internal platform for engineers who use coding agents, this is the conversation to have with your lead in the next two to four weeks. Settle the channel policy first: adopt the open Agent Skills format for portability, route installs through a private marketplace, and use managed settings to keep unreviewed mods off developer machines until someone has read their hooks. Require every internal skill to ship with its own verification against the live system, as universal-modder does with real-game testing and a pre-release check, and keep skills small, since SkillsBench found two or three focused modules beat comprehensive documentation. Treat stars and directory counts as attention, not quality, and measure your own pass rates instead. Then size MCP and orchestration hiring for a market where roles are visible and proven talent is hard to verify, and consider contractors for the first production servers.
Three public events fall inside or just after the next two weeks. LangChain's Interrupt agent conference reaches London on October 13 at Outernet. The Linux Foundation and the Agentic AI Foundation host AGNTCon and MCPCon North America on October 22 and 23 in San Jose, one of the foundation's two flagship events. GitHub Universe follows on October 28 and 29 at Fort Mason Center in San Francisco.
Registries can count what was installed, but only a loop that checks the work can say what was worth installing.
Sources:
- Make Claude Code mod any PC game you own: a hands-on guide to universal-modder - AI Frontier Post (https://aifrontierpost.com/articles/universal-modder-ai-game-modding-tutorial/)
- Universal Modder: A Practical Guide to AI Game Modding (https://huggingface.co/blog/a2aprotocol/universal-modder-a-practical-guide-to-ai-game-modd)
- Claude Code Mods Just Turned Agents Into a Platform - AgentConn Blog (https://agentconn.com/blog/claude-code-mods-app-store-moment-2026/)
- Mods overview - Claude Code Docs (https://code.claude.com/docs/en/plugins/mods/overview)
- Manage mods for your organization - Claude Code Docs (https://code.claude.com/docs/en/plugins/mods/admin)
- The Agent Skills Ecosystem in 2026: Who's Building, What's Working, and What's Next (https://agentman.ai/blog/agent-skills-ecosystem-report-2026)
- Claude Skills Directory: Browse 23,600+ Claude Code Skills - claudemarketplaces.com (https://claudemarketplaces.com/skills)
- Claude Code Plugin Marketplaces - claudemarketplaces.com (https://claudemarketplaces.com/marketplaces)
- GitHub - Louis-CFM/coucou (https://github.com/Louis-CFM/coucou)
- GitHub - storytold/photocraft (https://github.com/storytold/photocraft)
- arXiv 2610.03938: MLLMs Fail to Refuse when Using Tools Agentically (https://arxiv.org/abs/2610.03938)
- arXiv 2610.04053: The Cost of a Hop: Benchmarking NLIP and A2A (https://arxiv.org/html/2610.04053v1)
- arXiv 2610.04156: Trajectory-Derived Confidence for Reliable, Resource-Aware Clinical Text-to-SQL Agents (https://arxiv.org/abs/2610.04156)
- arXiv 2601.10338: Agent Skills in the Wild: An Empirical Study of Security Vulnerabilities at Scale (https://arxiv.org/abs/2601.10338)
- arXiv 2602.12670: SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks (https://arxiv.org/abs/2602.12670)
- Hiring AI Agent Developers: Skills, Sourcing, and Interview Strategies - daily.dev Recruiter (https://recruiter.daily.dev/resources/hiring-ai-agent-developers-skills-sourcing-interview-strategies/)
- MCP Developer Salary 2026: What to Pay (Global Ranges) - AY Automate (https://www.ayautomate.com/blog/mcp-developer-salary)
- The Agent Conference by LangChain, London (https://interrupt.langchain.com/london)
- Agentic AI Foundation Announces Global 2026 Events Program Anchored by AGNTCon and MCPCon North America and Europe - Linux Foundation (https://www.linuxfoundation.org/press/agentic-ai-foundation-announces-global-2026-events-program-anchored-by-agntcon-mcpcon-north-america-and-europe)
- GitHub Universe 2026 (https://reg.githubuniverse.com/flow/github/universe26/cfs/page/cfs-landing)
Disclaimer: This article is generated by GatiFlow Intelligence for informational purposes only. It does not constitute investment advice, recruitment recommendations, or legal guidance. All data is derived from public sources and AI analysis — verify independently before making decisions. Past trends do not guarantee future results.
Where this came from
Every Deep Dive starts from GatiFlow's own pipeline: 13 public developer sources, collected every six hours, with a confidence score and the evidence behind each signal. The same signals, filtered to the topics you follow, are a JSON API.
No credit card required.
Get the next one by email
One article every Saturday morning in your time zone. No account needed, and nothing else is sent to the address.
Double opt-in: you confirm by email first. What we store, and for how long, is in the privacy policy.
Tell me I am wrong
Corrections, the version of this you have lived through, or what you would like covered next. It reaches me directly and is never published. It is kept for two years so it can be read and answered; the privacy policy has the details.
0/2000