Lecture VII: Extending the Agent

AI-Assisted Programming for PhD Researchers

Dr. Tobias Vlćek

Helmut Schmidt University, August 2026

Why Extend

Out of the Box

Your agent already has three things:

  • The files in your project
  • A shell to run commands
  • Its training data: frozen, and often out of date

It lacks today’s docs, your library’s real API, your reference manager, and your recurring workflows.

Four Extension Mechanisms

This session tours all four, simplest to most powerful:

  • Skills: canned instructions it loads on demand
  • MCP servers: give the agent new tools
  • Subagents: delegate a subtask to a fresh context
  • Worktrees: run work in parallel without collisions

We start with the simplest, a plain markdown file, and build up to tools and parallel agents.

Skills

Recurring Instructions, Packaged

  • The lowest-effort extension: just a versioned markdown file
  • No install, no server, no keys: you write instructions and save them
  • The agent reads a skill only when the task matches its description
  • Keeps your context lean until the moment the knowledge is needed

Skill or AGENTS.md?

  • AGENTS.md: always-on rules applied to every interaction
  • A skill: on-demand expertise the agent loads only when the task matches its description
  • Reach for a skill when the instructions are for a specific recurring task, not every session
  • You already maintain an AGENTS.md from Labs 1 and 2; a skill sits beside it
  • Both are just versioned markdown you own

The Course Writing Skills

We ship adapted skills you will run on real text this afternoon:

  • academic-grammar, conciseness, hedging-language: grammar, filler, over-claiming
  • acronym-check, term-consistency, tense-consistency: the consistency trio
  • humanizer: strip AI-tell phrasing
  • literature-grounder: add grounded citations
  • Eight editors, each doing one job

Skills Compose

  • Small, single-purpose skills chain together over one piece of text
  • Run academic-grammarconcisenesshedging-language on a paragraph
  • Each does one job, so each diff is easy to review
  • A skill can even reach an external API with no server: reference-lookup queries OpenAlex for a DOI
  • The rule of thumb: skills before servers

Anatomy of a Skill

Frontmatter says when; the body says how:

---
name: figure-caption-checker
description: Check figure captions for completeness. Use when reviewing figures in a manuscript.
---

Check each figure caption against these rules:
- State what the figure shows, how it was measured, and n.
- Report units on every axis and every value.
- No interpretation in the caption: that belongs in the text.
- Flag any caption that breaks a rule and suggest a fix.

Write Your Own

  • Any workflow you explain to the agent twice deserves a skill
  • Keep them short: a page of instructions, not an essay
  • Version them with your project, like any other file
  • A stale skill misleads; prune it when the workflow changes
  • Vet a downloaded skill by reading it, not its star count. Flashy repos are often fake-starred, and a skill carries supply-chain risk like any dependency

Borrow a Framework: Superpowers

  • Superpowers (MIT, Jesse Vincent): open-source workflow skills; I use it daily
  • Not knowledge but process discipline: planning, test-driven development, debugging, verification
  • Two ship in the starter: systematic-debugging (“root cause before fixes”) and verification-before-completion (“evidence before claims”)
  • They encode Day 2’s rules as skills; reading them is the vetting exercise

Model Context Protocol

What MCP Is

  • An open standard for connecting agents to external systems
  • A server exposes tools: functions the agent can call
  • Servers exist for documentation lookup, GitHub, web search, databases, and more
  • Any MCP-aware host can use them: OpenCode, Claude Code, Zed
  • “USB for agents”: write the server once, plug it in anywhere

The Problem It Solves

  • A model’s knowledge of a library is frozen at its training snapshot
  • Every release after that, the gap widens: renamed arguments, dead functions
  • This is where invented function signatures come from
  • Same failure as tomorrow’s invented citations, one layer down

Wiring One Up

Context7 looks documentation up as you ask. Hosted, so a remote URL and no command to install. Add it to opencode.json, then restart:

{
  "mcp": {
    "context7": {
      "type": "remote",
      "url": "https://mcp.context7.com/mcp",
      "enabled": true
    }
  }
}

No login, no token. The agent sees the tools and decides when to call them.

Live: Ask for Docs

You type a plain request:

How do I write a parametrized test in pytest,
and what arguments does the decorator take?
use context7
  • The agent resolves “pytest” into a library ID
  • It pulls the current documentation for it
  • It answers from that text instead of from memory

Other Servers, Same Shape

  • GitHub’s hosted server: same remote entry, different URL
  • It lists issues and reads pull requests without cloning the repo
  • But it needs auth, and a token scopes what it can reach in your account
  • Context7 reads public docs, so the worst case is a bad answer. A write-scoped GitHub token is a different order of risk

MCP Is Also an Attack Surface

  • A local server runs code on your machine with your permissions
  • Its tool descriptions enter your context, a channel for prompt injection

Same lethal trifecta as this morning: install only servers you trust (Willison 2025).

Subagents

Delegation with a Clean Slate

  • Spawn a fresh-context agent for a self-contained subtask
  • Good for research or a review that would flood your main session
  • It does the messy work and returns only a summary
  • Your own context stays lean and focused

The Fresh-Eyes Reviewer

  • This morning’s Lab 4 second reviewer was this pattern, done by hand
  • A subagent automates it: spawn one to review the code just written
  • Fresh context means no attachment to the code it reviews
  • It never saw the reasoning that produced the bug, so it questions it

When (Not) to Delegate

  • Good: broad exploration, independent verification, bulk research
  • Bad: tiny tasks (the spin-up overhead outweighs the work)
  • Bad: anything needing your judgment partway through

A subagent returns a summary, not a transcript. Delegate only when the summary is what you actually need.

Git Worktrees

Two Agents, One Repo?

  • Two agents in one working tree is chaos: both edit the same files
  • Branches alone do not help: only one branch is checked out at a time
  • You need a second checkout, not just a second branch

Worktrees: Extra Checkouts

A worktree is a second folder backed by the same repository:

# create a sibling folder on a NEW branch (-b creates it)
git worktree add -b docs-branch ../proj-docs
# see every checkout attached to this repo
git worktree list
# clean up when done (the branch survives)
git worktree remove ../proj-docs

Or from Zed’s UI: the git: worktree picker creates, opens, and removes them.

Run a second agent in ../proj-docs safely: separate files.

Caveats

  • Worktrees isolate files, not databases, servers, ports, or caches
  • Two agents sharing one dev database still collide
  • Clean up with git worktree remove so stale folders do not pile up

Today: one parallel docs-improvement worktree while your main session keeps working. One stunt, not a habit.

Automation Outlook

Agents Without a Chair

Beyond interactive sessions, agents can run unattended:

  • Headless runs: opencode run "..." with no human in the loop
  • Scheduled jobs: nightly cleanups, dependency bumps
  • CI bots: reviewing pull requests or fixing issues automatically

Powerful, and it inherits every risk from this morning. Guardrails and checks come first. This is a pointer to the literature, not course material.

Lab 5

Your Mission

  • Run a writing skill on real prose and review the diff
  • Write your own skill for a workflow you repeat (optional)
  • Build a verified reading list with the reference-lookup skill: open every DOI yourself
  • Run two agents in parallel with a worktree, then merge them (demoed live)
  • Wire Context7 into opencode.json and check its answer against the real docs (optional)

Continue Your Journey

Next Up

Willison, Simon. 2025. “The Lethal Trifecta for AI Agents: Private Data, Untrusted Content, and External Communication.” In Simonwillison.net. https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/.