Lecture VII: Extending the Agent

AI-Assisted Programming for PhD Researchers

Author
Affiliation

Dr. Tobias Vlćek

Helmut Schmidt University, August 2026

Why Extend

Out of the Box

Your agent already has three things:

  • The files in your project
  • A shell to run commands
  • Its training data: frozen, and often out of date

. . .

It lacks today’s docs, your library’s real API, your reference manager, and your recurring workflows.

Four Extension Mechanisms

This session tours all four, simplest to most powerful:

  • Skills: canned instructions it loads on demand
  • MCP servers: give the agent new tools
  • Subagents: delegate a subtask to a fresh context
  • Worktrees: run work in parallel without collisions

We start with the simplest, a plain markdown file, and build up to tools and parallel agents.

Skills

Recurring Instructions, Packaged

  • The lowest-effort extension: just a versioned markdown file
  • No install, no server, no keys: you write instructions and save them
  • The agent reads a skill only when the task matches its description
  • Keeps your context lean until the moment the knowledge is needed

Skill or AGENTS.md?

  • AGENTS.md: always-on rules applied to every interaction
  • A skill: on-demand expertise the agent loads only when the task matches its description
  • Reach for a skill when the instructions are for a specific recurring task, not every session
  • You already maintain an AGENTS.md from Labs 1 and 2; a skill sits beside it
  • Both are just versioned markdown you own

The Course Writing Skills

We ship adapted skills you will run on real text this afternoon:

  • academic-grammar, conciseness, hedging-language: grammar, filler, over-claiming
  • acronym-check, term-consistency, tense-consistency: the consistency trio
  • humanizer: strip AI-tell phrasing
  • literature-grounder: add grounded citations
  • Eight editors, each doing one job

Skills Compose

  • Small, single-purpose skills chain together over one piece of text
  • Run academic-grammarconcisenesshedging-language on a paragraph
  • Each does one job, so each diff is easy to review
  • A skill can even reach an external API with no server: reference-lookup queries OpenAlex for a DOI
  • The rule of thumb: skills before servers

Anatomy of a Skill

Frontmatter says when; the body says how:

---
name: figure-caption-checker
description: Check figure captions for completeness. Use when reviewing figures in a manuscript.
---

Check each figure caption against these rules:
- State what the figure shows, how it was measured, and n.
- Report units on every axis and every value.
- No interpretation in the caption: that belongs in the text.
- Flag any caption that breaks a rule and suggest a fix.

Write Your Own

  • Any workflow you explain to the agent twice deserves a skill
  • Keep them short: a page of instructions, not an essay
  • Version them with your project, like any other file
  • A stale skill misleads; prune it when the workflow changes
  • Vet a downloaded skill by reading it, not its star count. Flashy repos are often fake-starred, and a skill carries supply-chain risk like any dependency

Borrow a Framework: Superpowers

  • Superpowers (MIT, Jesse Vincent): open-source workflow skills; I use it daily
  • Not knowledge but process discipline: planning, test-driven development, debugging, verification
  • Two ship in the starter: systematic-debugging (“root cause before fixes”) and verification-before-completion (“evidence before claims”)
  • They encode Day 2’s rules as skills; reading them is the vetting exercise

Model Context Protocol

What MCP Is

  • An open standard for connecting agents to external systems
  • A server exposes tools: functions the agent can call
  • Servers exist for documentation lookup, GitHub, web search, databases, and more
  • Any MCP-aware host can use them: OpenCode, Claude Code, Zed
  • “USB for agents”: write the server once, plug it in anywhere

The Problem It Solves

  • A model’s knowledge of a library is frozen at its training snapshot
  • Every release after that, the gap widens: renamed arguments, dead functions
  • This is where invented function signatures come from
  • Same failure as tomorrow’s invented citations, one layer down

Wiring One Up

Context7 looks documentation up as you ask. Hosted, so a remote URL and no command to install. Add it to opencode.json, then restart:

{
  "mcp": {
    "context7": {
      "type": "remote",
      "url": "https://mcp.context7.com/mcp",
      "enabled": true
    }
  }
}

No login, no token. The agent sees the tools and decides when to call them.

Live: Ask for Docs

You type a plain request:

How do I write a parametrized test in pytest,
and what arguments does the decorator take?
use context7

. . .

  • The agent resolves “pytest” into a library ID
  • It pulls the current documentation for it
  • It answers from that text instead of from memory

Other Servers, Same Shape

  • GitHub’s hosted server: same remote entry, different URL
  • It lists issues and reads pull requests without cloning the repo
  • But it needs auth, and a token scopes what it can reach in your account
  • Context7 reads public docs, so the worst case is a bad answer. A write-scoped GitHub token is a different order of risk

MCP Is Also an Attack Surface

  • A local server runs code on your machine with your permissions
  • Its tool descriptions enter your context, a channel for prompt injection

Same lethal trifecta as this morning: install only servers you trust (Willison 2025).

Subagents

Delegation with a Clean Slate

  • Spawn a fresh-context agent for a self-contained subtask
  • Good for research or a review that would flood your main session
  • It does the messy work and returns only a summary
  • Your own context stays lean and focused

The Fresh-Eyes Reviewer

  • This morning’s Lab 4 second reviewer was this pattern, done by hand
  • A subagent automates it: spawn one to review the code just written
  • Fresh context means no attachment to the code it reviews
  • It never saw the reasoning that produced the bug, so it questions it

When (Not) to Delegate

  • Good: broad exploration, independent verification, bulk research
  • Bad: tiny tasks (the spin-up overhead outweighs the work)
  • Bad: anything needing your judgment partway through

A subagent returns a summary, not a transcript. Delegate only when the summary is what you actually need.

Git Worktrees

Two Agents, One Repo?

  • Two agents in one working tree is chaos: both edit the same files
  • Branches alone do not help: only one branch is checked out at a time
  • You need a second checkout, not just a second branch

Worktrees: Extra Checkouts

A worktree is a second folder backed by the same repository:

# create a sibling folder on a NEW branch (-b creates it)
git worktree add -b docs-branch ../proj-docs
# see every checkout attached to this repo
git worktree list
# clean up when done (the branch survives)
git worktree remove ../proj-docs

Or from Zed’s UI: the git: worktree picker creates, opens, and removes them.

Run a second agent in ../proj-docs safely: separate files.

Caveats

  • Worktrees isolate files, not databases, servers, ports, or caches
  • Two agents sharing one dev database still collide
  • Clean up with git worktree remove so stale folders do not pile up

Today: one parallel docs-improvement worktree while your main session keeps working. One stunt, not a habit.

Automation Outlook

Agents Without a Chair

Beyond interactive sessions, agents can run unattended:

  • Headless runs: opencode run "..." with no human in the loop
  • Scheduled jobs: nightly cleanups, dependency bumps
  • CI bots: reviewing pull requests or fixing issues automatically

. . .

Powerful, and it inherits every risk from this morning. Guardrails and checks come first. This is a pointer to the literature, not course material.

Lab 5

Your Mission

  • Run a writing skill on real prose and review the diff
  • Write your own skill for a workflow you repeat (optional)
  • Build a verified reading list with the reference-lookup skill: open every DOI yourself
  • Run two agents in parallel with a worktree, then merge them (demoed live)
  • Wire Context7 into opencode.json and check its answer against the real docs (optional)

Continue Your Journey

Next Up

References

Willison, Simon. 2025. “The Lethal Trifecta for AI Agents: Private Data, Untrusted Content, and External Communication.” In Simonwillison.net. https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/.