Lecture IX: Judgment and Showing Your Work
AI-Assisted Programming for PhD Researchers
When Not to Use AI
The Question Is Not “Can It?”
- It usually can produce something. That was never the question
- The question is whether you should delegate this
- Now, in a thesis you will have to defend
- “It compiles” and “I can stand behind it” are different bars
Do Not Delegate
- What you are here to learn: the core methods of your field
- Architectural decisions you will live with in the publication
- Anything you cannot verify yourself
- Statistics you do not actually understand
The Verification Boundary
- Usable delegation ends where your ability to check ends
- Self-perception is a bad meter for this
. . .
- METR’s developers felt ~20% faster while being 19% slower (Becker et al. 2025)
- Growing that boundary outward is your development as a researcher: the goal is calibrated trust
Keeping Your Edge
- Deliberately code without AI sometimes: I often write sketches on paper
- Make explain-back a habit: say what the code does before you accept it
- Confidence in the tool predicts less critical thinking (Lee et al. 2025)
- The RCT, one last time: how you use it decides what you keep (Shen and Tamkin 2026)
Your Own Guardrails
- When the goal is learning, ask the agent for hints and next steps, not full solutions
- That is the exact design that erased the learning harm in the PNAS tutoring RCT (Bastani et al. 2025)
- The effect is specific to that guarded-tutor study: treat it as a pattern to copy, not a guarantee
Discussion
Two questions:
- Which task this week would you not delegate again?
- Have you ever accept something you couldn’t explain in your research?
Showing Your Work
This Afternoon
- I go around the room, everyone shows what they built
- No slides, no rehearsed talk, nothing to prepare over lunch
- Roughly ten minutes each, longer if it gets interesting
- Details: Showing Your Work
What to Show
- The thing running. The output, and the command that produces it
- How you worked. Open your spec and your
git loginstead of describing them - Your AI-use disclosure. Which tool, roughly how much came from the agent, where you overrode it
- What went wrong. Where you had to step in, what cost you an hour
What ís interesting to hear
- Where you caught the agent, and how
- Where you did not catch it until later
- Anything you would warn the next person about
Continue Your Journey
After the Course
- Keep the three habits: the loop, the checks, the review
- This website stays up. Come back to it
- Go deeper: course literature and references
- How this afternoon runs: Showing Your Work
- Questions later? tobiasvlcek.com
References
Bastani, Hamsa, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, and Rei Mariman. 2025. “Generative AI Without Guardrails Can Harm Learning: Evidence from High School Mathematics.” Proceedings of the National Academy of Sciences 122 (26): e2422633122. https://doi.org/10.1073/pnas.2422633122.
Becker, Joel, Nate Rush, Elizabeth Barnes, and David Rein. 2025. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv. https://doi.org/10.48550/ARXIV.2507.09089.
Lee, Hao-Ping (Hank), Advait Sarkar, Lev Tankelevitch, et al. 2025. “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers.” Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (Yokohama Japan), April, 1–22. https://doi.org/10.1145/3706598.3713778.
Shen, Judy Hanwen, and Alex Tamkin. 2026. How AI Impacts Skill Formation. arXiv. https://doi.org/10.48550/ARXIV.2601.20245.