More and more agents—so who makes the calls? Five technical takeaways from this week's coding-agent ecosystem
From Jev-style decision models and always-on agents to Agent Skills, harness research, multi-session monitoring, and hands-on Sonnet 5.5 reports
AI AgentSkillDev tools & CI/CDJevAlways-on agentsHarness16 min read
This week’s heatSummed community scores; bars show the last 8 weeks · click a name to open it on the map
Hottest
1,907
1,260
811
Rising
Needs data from the previous week (2026-W39). Shows from the next issue.
Cooling
Needs data from the previous week (2026-W39). Shows from the next issue.
This week the center of community discussion shifted clearly from "which model is strongest" to "the layer around the model." The most common new projects on GitHub were Jev-style fast-judgment models, open-source versions of always-on agents, and full Agent Skills bundles that ship with scripts and validators. Hugging Face's paper leaderboard, meanwhile, was full of harness and self-evolution research.
The arguments were remarkably consistent, too: agents are making more and more decisions—which ones should go to a cheap small model, which to a large model, and which must have a human sign off? How to scope permissions for always-on agents, what belongs in a Skill versus AGENTS.md, and whether self-evolution just amplifies mistakes are all facets of the same question.
This piece covers five topics in order. Each chapter stands on its own, and the end has things you can try right away.
Small models as a "fast judgment layer": what open alternatives exist beyond Jev?
The word that came up most in this week's new GitHub projects wasn't a big model—it was Jev. It's a decision API built for multiple-choice questions: you send in some state and a few questions, and it returns a structured answer in about 300 ms. The community has mapped out its three question types:
Yes/no: returns a probability from 0 to 1.
Multiple choice: picks only from the options you give, with a probability for each, and never invents new categories.
Rating: scores on a 2–10 scale you define.
The point isn't how clever the answers are—it's that they're fast, cheap, and calibrated. One HN commenter noted it can ingest 30k+ tokens in 500–800 ms and guessed it isn't a typical autoregressive model but something closer to a bidirectional (BERT-like) architecture. Others doubted that a 20–50k tokens/s prefill is possible on an ordinary LLM and suspected cache hits were inflating the numbers.
What are developers using it for?
disler's ten-levels-of-jev splits usage into ten levels. The first five use it as a "smart if" in code: single judgments, multi-class routing, weighted scoring, confidence thresholds, and sending simple requests to cheaper models. The last five embed it in a coding agent:
Guardrail hook: screen every tool call with Jev first.
Whether to compact context: decide after each turn whether it's time to compact.
Cheap file reads: decide whether a file is worth reading before stuffing it into context.
Parallel judgments over many files: ask about hundreds of files at once.
The most concrete implementation is jevgrep (★2,097), a CLI for coding agents: ask in plain language "which code does X," and it walks down the directory tree judging relevance, then returns a file list and line-numbered source snippets. On 10 SWE-bench tasks the author got the same 8/10 success rate, but coding-agent spend dropped about 28.6% (about 25.8% including Jev's own cost).
Comparing the open alternatives
Jev is a closed service, so several "run it yourself" versions appeared this week:
Aspect
Jev
jeff
jeeves
GLM-5.3-Flash rewrite
Form
Cloud API (closed)
0.8B open model
9B open model
Prompting trick on an open model
Base
Undisclosed
Qwen3.5-0.8B (also 2B, Gemma4-E2B)
Qwen3.5-9B + LoRA
GLM-5.3-Flash
Approach
Single inference, returns probabilities
One forward pass outputs per-option probabilities
Reason first, then choose; SFT + CISPO RL
Make the first output token the answer
Latency
~160–260 ms
~22 ms GPU, ~28 ms M4 Max, ~463 ms CPU
~3.3 s with reasoning, ~0.3 s without
~160–300 ms (depends on region)
Accuracy
Baseline
78.7% across 5 public test sets (Jev 83.0%)
0.889 on its test set (Jev 0.857)
Close to Jev, but worse calibration
License
Commercial service
Code MIT, weights Apache 2.0
MIT
Per model license
A few details worth noting:
jeff (★1,338) calls itself a "System 1 model." At just 1.7 GB, it can hot-load nine ~41 MB LoRA adapters at once (prompt-injection defense, support intent, tool selection, spam, legal clauses, and more) without restarting. The author puts it in front of Qwen 27B: confident cases are answered directly and uncertain ones go to the big model, for 95.3% overall accuracy at 38× the speed.
jeeves (★400) goes the other way: it reasons before answering, scoring 0.865 vs. Jev's 0.730 on a hard subset. It scores options with a "pointer head" and uses diffusion-style speculative decoding to speed up reasoning about 1.6×.
The GLM-5.3-Flash rewrite (138 HN points) trains nothing; prompt design alone makes the model answer in a single token, and it gains image input. But commenters pointed out that while the accuracy ranking keeps up, the probability distributions don't match, and an autoregressive architecture will eventually lose the real-time feel.
Community consensus: the value of these models is whether the probabilities are right, not just whether the answer is. If you use them for thresholds or routing, calibration matters more than accuracy.
Some also warned that Jev can be "wildly wrong" in untested situations, and that the flood of Jev posts this week looked suspiciously like marketing.
How to choose an always-on agent: Dots, Grok Bot, and the open-source versions
"Always-on agents" are AIs that keep running in the background without being asked: sorting email on a schedule, watching Slack, tracking a topic. This week Dots took 764 points and 644 comments on HN, and within days three open-source clones appeared on GitHub.
Hosted: Dots, Grok Bot, Manus
The Dots architecture as pieced together in the HN thread:
One cloud computer per agent: its own Linux environment and Chrome, able to browse, create files, and run tools, isolated between users.
Read-only by default: connected apps can be read but not used to send email, edit content, or operate the computer.
Automatic review: actions that need approval are intercepted and checked before running.
"Proactive research" runs in the background: using only read-only tools.
Commenters cared most about two things: Gmail and Outlook have no true read-only API keys, so authorizing often means granting full access; and background "proactive research" keeps burning tokens and compute. Plenty of people simply called it "hosted OpenClaw, plus single-vendor lock-in."
Grok Bot is a similar product that launched in August (HN discussion) with a different approach: it logs into your apps and operates their interfaces directly, keeping browser sessions signed in. The most-cited numbers in the thread were about security: one commenter cited a prompt-injection test claiming one model was compromised about 43% of the time and even the best about 5%—"but 5% is still a disaster for an always-on agent." Another hot topic was whether dangerous commands like git should be restricted to read-only.
Self-hosted, open source
Aspect
OpenDots
open-dot
feder-cr/dots
Stars
★1,875
★467
★2,565
Positioning
Multi-channel AI coworker (text, calls, Slack)
Personal always-on agent for macOS
Web agent with its own hard-to-block browser
Where it runs
Self-hosted; Docker split into app and browser containers
Local Mac app, starts a background server
Local, installed with uvx
Model
Set in .env
OpenAI Responses API (plus a small model to review rules and a realtime model for voice)
Any model via OpenRouter
Integrations
CopilotKit, AG-UI, Slack Channels SDK
Composio (1,500+ apps, OAuth)
Exposes browser capabilities to Claude Code and others as an MCP server
Natural-language "allow / ask first / deny" rules, each checked by a small model
Nothing special
License
MIT
Not stated
MIT
Key implementation points:
open-dot has the most complete permission design: reads can be automatic, writes always prompt for approval; passwords are encrypted with AES-256-GCM with the key in the macOS Keychain, so the agent never sees plaintext.
feder-cr/dots is really a browser layer: it patches the Firefox engine in C++, handling fingerprints, mouse paths, and per-keystroke typing inside the engine, then wraps it as an MCP server. It's meant to pair with other agent frameworks rather than serve as an always-on agent on its own.
OpenDots leans toward teams: Slack workspace and user allowlists, with credentials kept only on the server.
One table to sum it up
Aspect
Dots
Grok Bot
Self-hosted (OpenDots / open-dot / OpenClaw)
Where it runs
Vendor cloud sandbox
Vendor cloud, logged into your apps' interfaces
Your own computer or server
Default permissions
Read-only, writes need review
Gets account login access
You set them; varies widely by project
Model
Vendor-locked
Vendor-locked
Swappable (OpenRouter, local models)
Cost
Pro plan from $100/month; not yet available in the EU/UK
Tied to a subscription plan
Only model API costs
Main risk
Vendor lock-in, background token burn
Prompt injection, overly broad UI control
You maintain it; updates often break things
Community consensus: the hard part of always-on agents isn't what they can do but how read and write permissions are separated. Until APIs offer read-only options, self-hosted versions actually make least privilege easier.
How far have Agent Skills come? From AGENTS.md to full toolkits
Since 9/25, 2,232 new GitHub repos have been tagged with the agent-skills topic. The top-starred ones are no longer a single SKILL.md with a few prompts but complete packages of prompts + reference material + executable scripts + validators.
What this week's top Skills have in common
Aspect
universal-modder
logo-design-skill
yomiyasu
Stars
★2,371
★1,603
★1,251
Purpose
Lets a coding agent mod any PC game
Systematic logo design
Turns AI-written Japanese into natural Japanese
Reference material
Guides for 12 game engines
1,400+ categorized SVG logos
Lists of AI-sounding patterns, writing guidelines by field
Scripts/tools
um CLI with 15+ commands; fal MCP for art and audio
Launches the game, screenshots, records video, backs up saves
Checks at 16/32/96 px and in monochrome
Quantifies "AI flavor": bold density, list ratio, repeated sentence endings
Cross-agent support
Symlinks into .claude, .gemini, .cursor, .codex at once
Claude Code plugin, zip upload, generic Skill format
Claude plugin, npx skills add
Three trends stand out:
Checks written as code: the logo skill catches "skew angles off by 0.3–3 degrees," and yomiyasu caps bold text at 2 per thousand characters and lists at 15%. They no longer rely on the model to judge quality on its own.
Workflows with explicit checkpoints: universal-modder defines a 10-step loop (search the knowledge base → recon → set up a safe environment → read the source → build a small piece first → generate assets → verify in the real game → record → package → write notes for the next agent). The last step, "leave notes for the next agent," effectively saves experience back into the Skill.
One Skill, many agents: symlinks or npx skills let Claude Code, Codex, Gemini CLI, and Cursor all read the same copy.
Is AGENTS.md useful?
Skills load on demand; AGENTS.md/CLAUDE.md load every time. A research roundup reposted on HN this week (originally published 9/23) offered plenty of numbers:
Hand-written AGENTS.md raised success rates by 2.4 points; auto-generated ones lowered them by 2 points. Neither was significant, but inference cost rose about 20%.
One evaluation found that putting a compressed docs index in AGENTS.md passed 100% of the time; using a Skill passed only 53%, because the Skill was actually invoked only 44% of the time.
An analysis of 7,310 rules across 83 open-source projects found that keeping rules updated raised compliance from 49% to 72%.
Bottom line: knowledge the model doesn't have and needs every time goes in AGENTS.md; bulky material used only for specific tasks becomes a Skill.
Too many config files?
Having CLAUDE.md, AGENTS.md, .cursor/rules, and .mcp.json in one project is now common. Two approaches surfaced this week:
etymon: manage every config with one etymon.toml plus a lockfile, then etymon sync converts it into the formats of 18 tools including Claude Code, Codex, Cursor, and Gemini. It's very early with few stars, but the idea is worth a look.
Breadcrumb (44 HN points): attacks the context problem from the other end, recording your Mac's screen, meeting transcripts, and AI chats, encrypting them locally with SQLCipher, and letting agents query them through 30+ MCP tools. Some worried 30 tools would flood the context; the author replied that good clients lazy-load tool definitions.
Is the harness the new battleground? Agent frameworks and self-evolution research
"Harness" means the layer wrapped around the model: how tools are wired up, how context is managed, how actions are verified, how failures are retried. On Hugging Face's paper leaderboard this week, at least five titles mention harnesses directly; and the app with the most tokens on OpenRouter this week wasn't a chat interface but Hermes Agent (2.01T tokens), followed by Claude Code, Kilo Code, Cline, and Codex—all harnesses.
Treats "model + harness" as a composable unit: a lead agent splits tasks, dispatches them to specialist agents, and turns experience into reusable workflows (Skill Forge)
Among HF's most-upvoted this week; beats existing agent systems on long tasks
A "builder" agent builds a better environment for a "target" agent, learning principles for when help is needed and what to give
8.95 points above no Skill; 12.02 above handing the Skill to the target agent directly
Mid-Harness has a very practical finding: whether extra sampling helps depends on how strong the verifier is. With a weak verifier, 8 samples are about as good as 1; with a strong one, the same small model's success rate climbs sharply. It's the same idea as chapter one's "cheap judgment layer."
Self-evolution: agents that improve themselves
Another group of papers has agents improve using their own experience:
RSIGame: automated game development with two loops. The inner loop runs "explore → diagnose → fix" and accumulates an ever-growing test list; the outer loop monitors overall quality, keeps the best version, and detects stalls. Successful experience is then trained back into the model. On 140 Godot/Phaser tasks, a 27B open model generated 11× fewer tokens after training.
Self-Evolving Coding Agents: writes robot tasks as code (world state is code, policy is code) and trains the model on verified execution traces. Overall RoboCasa365 success rose from 56.6% to 60.8%.
hypoarena (★558): a "generate → debate → evolve" loop for scientific hypotheses, ranked by Elo tournament, with a pure NumPy core that runs on CPU.
The self-evolution trap: cheating together
False Frontiers points out an easily missed problem: when a question-setting agent and an answering agent train each other in a closed loop, they gradually agree on the same mistakes. Internal scores look better and better while external accuracy doesn't improve.
Aspect
Rate of shared errors
No mitigation
6.1%–8.8%
Multi-sample verification (MSV)
5.7%–7.2%
CrossFit (group A's questions get feedback only from agents trained on group B's data)
3.0%–3.7%
Excluding the original data entirely from feedback
~0.1%–0.4%
Community consensus: self-evolution needs checks from outside the loop, or it just amplifies its own mistakes.
Running several coding agents at once: how do you manage them? Plus hands-on Sonnet 5.5 reports
The always-on agents, Skills, and harnesses in the previous chapters all run into the same everyday problem: you have three or four Claude Code or Codex sessions open—which one is waiting for you to click "Allow," which has finished, and which is stuck?
Monitoring tools: getting status out through hooks
This week's most-starred project was coucou (★3,122), a little character that lives in your Mac's notch (or at the top of the screen on Windows/Linux) and keeps an eye on all your coding agents. Its approach is worth borrowing:
Hooks, not guesswork: it installs hooks in Claude Code, Gemini CLI, and Antigravity, and a tiny nb-hook script forwards events to the app—over a Unix socket on Mac/Linux and a named pipe on Windows.
Handle permission requests right in the notch: "Allow/Deny" buttons pop up (Claude Code also gets "Always allow"), so you don't have to switch back to the terminal.
Custom agents supported: add a coucou_agent tag to the hook payload and your own agents can plug in.
Cross-platform: native Swift 6 on Mac, Tauri 2 on Windows/Linux. Code is MIT; the character art is copyrighted separately.
Another direction is agent-office (★546): a cartoon 3D office where a team "hires" Claude Code workers, with shared live terminals, voice chat, and GitHub issue and PR tracking. It's built for collaboration, not personal use.
What they share: permission requests are where agents most often get stuck, so the point of a monitoring tool is letting you approve quickly from anywhere, not drawing a pretty dashboard.
Sonnet 5.5 and Opus 5.5: what do hands-on reports say?
The hottest model topic this week was Sonnet 5.5 (883 HN points, 614 comments), which also jumped straight to #1 in OpenRouter token usage this week (545B). Coding benchmarks reposted by the community:
Aspect
Sonnet 5.5
Opus 5.5
Terminal-Bench
70.6%
66.4%
FrontierCode
52.1%
54.4%
CursorBench
55.5%
57.8%
Speed
Faster
Slower
Cache read price
Same as Opus
—
Real-world experience from the comments:
The cheaper one wins on terminal tasks: Sonnet beats Opus by 4 points on Terminal-Bench, and many plan to switch their day-to-day coding to Sonnet.
No savings from caching: cache reads cost the same as Opus, so workflows that lean on prompt caching don't save much.
Security filtering is too strict: some got blocked just reading old code or doing legitimate security research, with automatic fallback to older models, so they moved to GLM and DeepSeek on OpenRouter. Others said GLM-5.3 Flash performs well on agentic coding at about 1/20 the price.
Another discussion, "Prompting Claude Opus 5.5" (207 HN points), collected a few prompting tips:
Don't just say "avoid looking AI-generated"—name the specific design patterns you don't want, ideally with reference sites.
For long runs of tool calls, ask it to report progress regularly, or users just see silence.
Some complained answers are still long with the point buried at the end; another said "every 3–6 months you have to relearn how to prompt, like relearning to drive every quarter."
What to try now
Add a cheap judgment layer to your agent workflow: use jeff (0.8B, runs locally) or a Jev-style API to pre-screen questions like "is this file worth reading" or "is this tool call safe," and send only the uncertain ones to a big model.
Start always-on agents read-only: whether you use Dots or self-host OpenDots/open-dot, enable reads only and require approval for every write; for finer control, borrow open-dot's "allow / ask first / deny" rule design.
Sort out your AGENTS.md and Skills: knowledge the model doesn't have and needs every time goes in AGENTS.md; bulky workflows become Skills, with their checks written as scripts. If you run several sessions at once, install a hook-based monitor like coucou.
Quiz
In the Mid-Harness experiments, what mainly decides whether "sampling several candidate actions per step" helps?
With a weak verifier, 8 samples are about as good as 1; with a strong one, a 9B model on TerminalBench-Lite goes from 50.0% to 68.0%.
According to this week's research roundup, what kind of knowledge belongs in AGENTS.md?
Bulky workflows suit Skills. AGENTS.md loads every time, and auto-generated versions actually lowered success by 2 points while raising inference cost about 20%.
What problem does False Frontiers say self-evolving agents are most prone to?
In a closed loop, internal scores keep improving while external accuracy doesn't; excluding the original data from feedback brings shared errors down to about 0.1%–0.4%.
One of the few posts with published AGENTS.md-vs-skills experiment data. Note Vercel also builds skills tooling; weigh it against the pushback in the comments.
Each week one AI technology topic, in depth: AI agents, skills, dev tools and CI/CD, OpenClaw, companion robots, and enterprise adoption. Only tools, open-source projects, implementation, architecture, and hands-on tests. No opinion pieces, policy, or business news.
Where the data comes from
Community discussions and independent evaluations only: Hacker News, Reddit, GitHub, Hugging Face, Medium, OpenRouter, Lobsters, dev.to. No official press releases, and only content from the past 7 days.
How it’s made
Saturday night, public APIs surface the week’s hot keywords. Once topics are picked, the article is written early Sunday, with every source listed at the end.
How the tech map grows
Every discussion an issue cites goes into the timeline, and relationships between technologies go into the graph. A new keyword that makes the screening list 3 weeks in a row gets promoted to a tracked technology.
How heat is scored
HN points + Reddit score + Hugging Face votes + GitHub stars ÷ 10. When one discussion covers several technologies, each one gets the full score.
Languages
Chinese is the original. The English version is translated from it by Claude; numbers and links are identical in both.