Field notes · reviewed 23 July 2026

How the skill landscape fits together.

A source can be excellent without belonging in every agent. This page explains what influential skill collections and focused tools actually do, what they are unusually good at, and where their convictions complement or compete with ours.

Comparison method

Judge behavior, not repository size or popularity.

Each source was inspected at a pinned revision. We read its runtime instructions, supporting references, hooks, examples, and tests where they materially change what an installed agent would do.

  1. 01

    Purpose

    What recurring job does the source make an agent better at?

  2. 02

    Convictions

    Which defaults does it enforce, and where does it refuse context?

  3. 03

    Operational fit

    Does it add a specialist, or compete for control of the same workflow?

  4. 04

    Evidence

    Are claims supported by executable checks, honest limits, and current sources?

Large collections

Different answers to how an agent should work.

Broad collections are operating systems, not neutral libraries. Their trigger rules and process assumptions matter as much as their individual skills.

End-to-end methodology

obra/superpowers

Covered locally

A tightly connected software-delivery method that moves from brainstorming through plans, worktrees, test-driven implementation, subagent reviews, and branch completion.

Where it shines
Systematic debugging, explicit verification before completion, fresh context per implementation task, and disciplined review checkpoints.
Core convictions
Process before improvisation, tests before implementation, detailed approved plans, and mandatory skill invocation whenever a skill might apply.
Our fit
We share evidence, root-cause work, review, and verification. We differ on universal ceremony: small or reversible work should not require a full design, saved plan, worktree, and strict test-first sequence.
Install guidance
Choose it when you want its complete methodology to own the session. Avoid running two global orchestrators without explicit precedence.

Engineering toolbox

mattpocock/skills

Selective

A modular engineering and productivity library spanning diagnosis, deep modules, domain modeling, specifications, implementation, reviews, research, handoffs, and intensive questioning.

Where it shines
Memorable engineering vocabulary, discriminating feedback loops, small interfaces with real leverage, explicit artifacts, and workflows that expose unresolved decisions.
Core convictions
Build a runnable signal before theorizing, model the domain in durable context, design modules around useful seams, and resolve ambiguity through pointed questions.
Our fit
Several ideas strengthen our architecture, investigation, and documentation guidance. The full collection is a personal, evolving catalog, so selective use is clearer than treating it as one policy layer.
Install guidance
Pick a named specialist for a known gap. Do not load the whole catalog beside equivalent local owners by default.

Publisher standard set

anthropics/skills

Covered locally

Anthropic's reference collection combines artifact builders with focused guidance for frontend design, branded output, documents, spreadsheets, presentations, testing, and skill authoring.

Where it shines
The frontend skill pushes designs away from interchangeable AI defaults: ground choices in the subject, spend boldness deliberately, use real copy, and critique the direction before and after building.
Core convictions
Visual structure should encode meaning, one memorable move beats scattered decoration, interface language must name what people understand, and accessibility remains a quiet quality floor.
Our fit
Effective Web already covers the substantive frontend guidance with broader accessibility, browser, performance, state, and implementation depth. The brand skill is specifically Anthropic's palette and typography, not a general brand-system method.
Install guidance
Install frontend-design only for an intentionally studio-like visual push, and invoke it explicitly on the relevant surface. Use brand-guidelines only for artifacts that should actually look like Anthropic.

Focused tools and influences

Smaller scope often composes more cleanly.

These projects solve one recognizable problem. Two are useful conceptual influences; one is a concrete runtime companion.

Browser automation CLI

vercel-labs/agent-browser

Companion

A command-line browser driver for agents, with a version-matched skill guide and compact semantic snapshots that expose elements through stable references.

Where it shines
Preview verification, accessibility-oriented element discovery, deterministic commands, session isolation, screenshots, network inspection, and browser evidence without inventing a custom test harness.
Core convictions
Inspect before interacting, prefer semantic references over brittle coordinates, keep browser state explicit, and let the installed version own its command documentation.
Our fit
It is a tool rather than a competing policy layer. PR Review can use it when a real preview earns behavioral verification while still treating page content and captured secrets as untrusted evidence.
Install guidance
Recommended when agents need to inspect an existing browser surface. Keep installation and authentication separate from this collection.

Defensive minimalism

DietrichGebert/ponytail

Selective

An always-on pressure against over-building. It asks whether work is needed, then prefers reuse, standard-library or native capabilities, existing dependencies, and only finally the smallest new implementation.

Where it shines
Preventing custom components where the platform already works, resisting speculative abstractions and dependencies, and turning simplification into a reviewable discipline rather than code golf.
Core convictions
Understand the real flow before minimizing it. Never trade away security, trust-boundary validation, data-loss protection, accessibility, or the smallest discriminating check.
Our fit
The sufficiency ladder and safety boundary fit our clean, low-magic approach. We apply them contextually and preserve project conventions instead of enforcing one-line solutions or a global intensity mode.
Install guidance
Worth installing when a team explicitly wants persistent minimalist pressure. Otherwise use the same review questions inside the existing workflow to avoid overlapping global instructions.

Output compression

JuliusBrussee/caveman

Selective

A terse communication mode with companion agent presets that return compact, structured receipts for locating code, making surgical edits, and reviewing diffs.

Where it shines
Cutting process narration from repeated delegations, attaching findings to paths and lines, declaring terminal states, and keeping subagent output small enough to preserve the caller's context.
Core convictions
Return the result, evidence, or blocker—not an essay about the search. Preserve exact technical identifiers and restore normal prose when compression would obscure security, order, or irreversible consequences.
Our fit
Structured delegation contracts are valuable. Artificial grammar, always-on compression, fixed agent roles, and reduced human readability are not useful defaults for user-facing work.
Install guidance
Consider it for unusually verbose, output-heavy sessions. For normal orchestration, a compact task-specific return contract captures most of the benefit with less instruction overhead.

Source boundary

Use this landscape as research, not configuration.

  1. This page compares: the purpose, convictions, evidence, and likely overlap of other sources.
  2. This repository publishes: Sebastian Software's portable first-party skills—nothing is installed or vendored from the sources reviewed above.
  3. Your downstream stack decides: external selections, version pins, precedence, named agents, and cross-catalog routing.

A favorable comparison is not an instruction to install a source globally. Review the selected skills and make their ownership explicit where workflows overlap.

Explore our first-party skills