Lecture IX: Judgment and Showing Your Work

AI-Assisted Programming for PhD Researchers

Dr. Tobias Vlćek

Helmut Schmidt University, August 2026

When Not to Use AI

The Question Is Not “Can It?”

  • It usually can produce something. That was never the question
  • The question is whether you should delegate this
  • Now, in a thesis you will have to defend
  • “It compiles” and “I can stand behind it” are different bars

Do Not Delegate

  • What you are here to learn: the core methods of your field
  • Architectural decisions you will live with in the publication
  • Anything you cannot verify yourself
  • Statistics you do not actually understand

The Verification Boundary

  • Usable delegation ends where your ability to check ends
  • Self-perception is a bad meter for this
  • METR’s developers felt ~20% faster while being 19% slower (Becker et al. 2025)
  • Growing that boundary outward is your development as a researcher: the goal is calibrated trust

Keeping Your Edge

  • Deliberately code without AI sometimes: I often write sketches on paper
  • Make explain-back a habit: say what the code does before you accept it
  • Confidence in the tool predicts less critical thinking (Lee et al. 2025)
  • The RCT, one last time: how you use it decides what you keep (Shen and Tamkin 2026)

Your Own Guardrails

  • When the goal is learning, ask the agent for hints and next steps, not full solutions
  • That is the exact design that erased the learning harm in the PNAS tutoring RCT (Bastani et al. 2025)
  • The effect is specific to that guarded-tutor study: treat it as a pattern to copy, not a guarantee

Discussion

Two questions:

  • Which task this week would you not delegate again?
  • Have you ever accept something you couldn’t explain in your research?

Showing Your Work

This Afternoon

  • I go around the room, everyone shows what they built
  • No slides, no rehearsed talk, nothing to prepare over lunch
  • Roughly ten minutes each, longer if it gets interesting
  • Details: Showing Your Work

What to Show

  • The thing running. The output, and the command that produces it
  • How you worked. Open your spec and your git log instead of describing them
  • Your AI-use disclosure. Which tool, roughly how much came from the agent, where you overrode it
  • What went wrong. Where you had to step in, what cost you an hour

What ís interesting to hear

  • Where you caught the agent, and how
  • Where you did not catch it until later
  • Anything you would warn the next person about

Continue Your Journey

After the Course

Bastani, Hamsa, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, and Rei Mariman. 2025. “Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics.” Proceedings of the National Academy of Sciences 122 (26): e2422633122. https://doi.org/10.1073/pnas.2422633122.
Becker, Joel, Nate Rush, Elizabeth Barnes, and David Rein. 2025. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv. https://doi.org/10.48550/ARXIV.2507.09089.
Lee, Hao-Ping (Hank), Advait Sarkar, Lev Tankelevitch, et al. 2025. “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers.” Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (Yokohama Japan), April, 1–22. https://doi.org/10.1145/3706598.3713778.
Shen, Judy Hanwen, and Alex Tamkin. 2026. How AI Impacts Skill Formation. arXiv. https://doi.org/10.48550/ARXIV.2601.20245.