Human-in-the-Loop Task Manager for AI Agents
Back to Blog
September 24, 2026 | AgentRQ Team

You Fixed the Playbook on Your Laptop. Nine Other Agents Never Heard.

Your release playbook had a bug in it. Step four said run the migration, then flip the feature flag, and it should always have been the other way around — the flag has to be off before the column moves or the read path sees a table that is half there.

You found it the hard way on Tuesday, and you fixed it properly. You opened ~/.claude/skills/release/SKILL.md, swapped the two steps, and added a sentence explaining why the order matters so nobody flips it back. It took ninety seconds.

On Wednesday, the agent on the build box ran the release. Migration first. Flag second.

It was not being careless. It was following the playbook — its copy of the playbook, the one in its own home directory, the one you never touched because you were not sitting at that machine. Somewhere on your laptop the correct version was sitting in a file, and it may as well have been in a notebook in a drawer.

A Skill Is a Playbook. It Was Also a Directory on One Machine.

A skill is a small idea that turns out to carry a lot of weight. It is a SKILL.md that says what the skill is for and when to use it, plus whatever reference files, prompts and scripts that file points to. An agent reads the description, decides the task in front of it matches, loads the playbook, and follows it. Claude Code works this way. Antigravity works this way. The format is good, and it is the same format in both.

The format was never the problem. Delivery was.

A skill lived in a directory in one agent's home, on one machine, and every way of getting it somewhere else was a copy. That breaks in three directions at once.

It is per machine. The agents worth running are not all on your laptop. One is on the office box because it has the database. One is on a VM because the job takes four hours. One wakes up on a schedule at 3am when nobody is watching, which is exactly the agent you least want running a stale playbook. Each has its own copy of the directory, each copy was correct on the day it was made, and none of them hear about Tuesday.

It is per tool. The same playbook has to exist under whatever directory each harness reads. Run Claude Code on one task and Antigravity on the next because it is better at that job, and the knowledge does not travel with you — not because the knowledge changed, but because it was filed in a place the other tool never looks.

Nothing owns the current version. This is the one that quietly costs the most. With five copies of a skill and no source of truth, "which one is right?" has no answer short of diffing all five. So people stop fixing skills in place. They paste the correction into the task instead, every time, which is exactly the work a skill was supposed to remove.

Skills Moved to the Workspace

In v0.8.0, skills live in AgentRQ.

Every workspace has its own skills. They are stored on the AgentRQ server — under a configurable local directory (AGENTRQ_STORAGE_DIR), or in any S3-compatible bucket (AGENTRQ_SKILLS_STORAGE=s3), grouped by workspace and skill. Only the metadata goes in the database. No agent's home directory is involved anywhere.

It is the same SKILL.md format, deliberately, so the skills you already have work unchanged:

markdown
---
name: pr-reviewer
description: Use when reviewing a pull request, before leaving any comments.
---

# Reviewing a pull request

1. Read the description, then `references/checklist.md`.
2. Run the tests before commenting on style.

description is the only required key, because a description is what an agent reads to decide whether this skill is the one. name is optional and defaults to the directory name. Any other frontmatter you were already using is kept as-is.

The Skills tab in an AgentRQ workspace's settings, listing skills imported from obra/superpowers with their descriptions, sizes and source commit, above a field for importing more from GitHub

Those fifteen are obra/superpowers, imported by pasting the repository link — every skill in it fits the limits, so all of them came in, each stamped with the commit it came from.

The Agent Does the Fetching

The interesting half is that nothing pushes skills to anyone. Agents come and get them.

Every agent connected to a workspace gets four new tools on that workspace's MCP server:

Tool What it does
searchSkills(q?, limit?, offset?) The skills this workspace can use, each with its description and skill:// URI. Contents are not included
loadSkill(uri) One file, as stored. A SKILL.md comes back with the URIs of the skill's other files
saveSkill(uri, content) Writes one file of one of this workspace's own skills. Writing SKILL.md creates or updates the skill
deleteSkill(uri) Deletes a skill, or one file of one

Files are addressed by URI. skill://systematic-debugging/SKILL.md is one file; skill://systematic-debugging on its own means that skill's SKILL.md.

The instructions every connecting agent receives now end with a step about this, in the same place the memory step lives: call searchSkills at the start of a task, load the SKILL.md of any skill whose description matches, then follow it — and load a skill's other files only when its SKILL.md points you to them.

That last clause is the design, and the comment above the tools in the source says why plainly: the tools disclose progressively, because loading every skill up front would spend the context on the ones the task does not need. A search returns names and descriptions. A load returns one file. The checklist a skill mentions in step one is fetched in step one, not before.

Shared, Not Copied

A workspace can share one of its skills with another workspace in the same account.

The word doing the work there is shared, not copied. The other workspace does not receive a duplicate to keep in sync; it reads the same skill, live. Fix a step in the owning workspace and every workspace it is shared with is already correct — not after a sync, not on next pull, immediately, because there was only ever one of them.

The ownership rule follows from that and stays simple: only the workspace that owns a skill can change or delete it. Everywhere else it is read-only. That is the difference between a shared playbook and five playbooks that agree for now.

A skill open in the Skills tab, showing the SKILL.md for releasing-payments with its steps, clickable references to preflight.md and rollback.md, a size readout of 1% of the 96 KB limit, and a share control showing it is shared with the billing-web workspace

That skill was not typed into a form. An agent wrote it with saveSkill, and the same panel reads it back, follows its references, and shares it with another workspace.

Bring the Skills You Already Have

Nobody wants to retype their skills into a new box, so there is a one-click import. In a workspace's Settings → Skills, paste a public GitHub link — a repository, or one skill's folder — and choose Import.

If the repository carries a plugin manifest, .agentrq/plugin.json or Muse's .muse-plugin/plugin.json, its skills list decides what to bring. Otherwise every directory holding a SKILL.md is a skill. The import keeps every Markdown file in a skill's folder, plus any other file the skill actually references — by path, by relative path, or through a folder written with a trailing slash — and leaves the rest.

Two details make this survive contact with real repositories.

Very large repositories are not downloaded whole. Past 20 MB compressed or 64 MB unpacked — garrytan/gstack, for example — the tab lists the repository's skills instead, each with its SKILL.md size, greying out any that exceed a limit. Tick the ones you want, choose Import selected, and only those files are fetched.

The import report says what did not come in, and why. A SKILL.md over the size limit, which skips that whole skill. A binary, symlink or hidden file. A file nothing references. A name the workspace is already using. The failure mode an import tool must not have is quietly bringing in 90% of something and letting you find out later.

The import picker for garrytan/gstack, which is too large to import whole, listing 60 skills with checkboxes and sizes; design-review is greyed out with the reason SKILL.md is 131838 bytes, the limit is 96 KiB. Below it, a skill shared from the payments-api workspace carries a SHARED FROM PAYMENTS-API badge

garrytan/gstack is the example in the docs, and this is what it looks like: sixty skills offered, one of them greyed out with the exact number that disqualified it. Producing that list downloads no skill files at all — only the repository's file tree, which is where the sizes come from, and a plugin manifest if there is one. The files arrive after you have ticked something.

What This Unlocks

A playbook gets one owner and one current version. This is the whole reason the Tuesday fix did not reach Wednesday's agent: the playbook was correct somewhere and nobody could say where. Centralising the storage collapses that question. There is one skill, it has an owning workspace, and editing it is the deployment — there is no second step where the change reaches the machines, because the machines never had it in the first place. They ask each time. A correction you make during a task is worth making, because you know it will hold.

The number of playbooks you can keep stops being bounded by context. The usual way to give agents standing procedures is to write them into CLAUDE.md or a mission description, which means every task pays for every procedure whether or not it is relevant. That ceiling arrives fast — a dozen good playbooks and you are spending real context on instructions for work this task will never do, so you start deleting useful things to make room. Progressive disclosure removes the tradeoff. A skill the task does not match costs its one-line description and nothing else; only the matching one is loaded, and only the parts of it that get pointed to. You can keep the playbook for the quarterly data migration next to the playbook for reviewing a PR, and the PR review does not pay for the migration.

saveSkill closes the loop, and that is the part that compounds. An agent that works out a better way to do something — the flaky test that needs a retry, the deploy gate everyone forgets, the API that returns 200 on failure — can write it back into the workspace's own skill instead of discovering it again next week. Combine that with live sharing and the improvement is not local: the next agent in every workspace that reads that skill starts from it, on whatever machine, in whatever harness. This is workspace memory one level up. Memory is what happened here; a skill is how this kind of work is done. Getting both out of an agent's home directory is what turns a fleet of agents that each start from zero into one that gets better at your work over time.

What It Deliberately Isn't

The edges, stated plainly, because a playbook system that oversells itself is worse than none:

  • → The limits refuse rather than truncate. A SKILL.md can be 96 KiB, every other file 64 KiB of UTF-8 text, and a skill can hold 64 files including its SKILL.md. Names are lowercase slugs up to 64 characters, descriptions up to 1024. Go over and it is rejected with a message saying what to change — a skill silently cut in half is a playbook with the last step missing, which is worse than no playbook.
  • → Sharing does not cross accounts. A skill can be shared with other workspaces in the same account, and that is the whole scope. There is no cross-account sharing and no public directory.
  • → A shared skill is read-only where it is borrowed. Agents there can load it; they cannot saveSkill over it. Improvements go through the owner.
  • → The Skills tab lists skills, not files. Opening one shows its SKILL.md, and you reach its other files by following the references inside it, with a breadcrumb showing where you are. It is the same progressive path the agent takes, which makes it easy to see what the agent will actually see.
  • → Names must be unique across everything a workspace can use, its own skills and the ones shared into it. An import that collides tells you, rather than shadowing something.

Where to Look

It shipped in v0.8.0, across a stack of PRs: #664 for the storage, validation, REST API and GitHub import, #665 for the MCP tools, #666 for the Skills tab, #668 for search, #673 for S3 storage and #674 for picking skills out of a large repository. The user guide is docs/SKILLS.md.

One more thing in the same release, small but worth knowing if you use Claude Code: the workspace Setup tab now pre-approves AgentRQ's tools with a single wildcard rule, mcp__agentrq-<workspace id>__*, instead of one line per tool (#682). Tools added to the server later — the four above, for instance — are covered without you editing anything.

If you have an agent connected to a workspace, it has searchSkills now. The useful first move is not to import anything: ask it what skills the workspace has, then give it one.

---

AgentRQ is currently in public beta. Join our GitHub community to help shape the future of human-agent collaboration.

Start Free