Session 1 — an introduction to Generative AI
Leibniz Institute for Financial Research SAFE
6 October 2026
Workshop synthesis, drawing on Mollick (2026), Velikov (2026) and Korinek (2025).
Session 2. Hands-on with your own materials: a referee report on your draft, secure data, writing, the paper–code link.
Slides, skills and guides for both days: claesbackman.com/ai_workshop.
A token is a word or a piece of a word, about four characters of English. That is the whole training objective: given the text so far, predict the next token. Press the button.
It does not look facts up. It generates text that is plausible given what it has seen.
It predicts what an answer looks like. The arithmetic is right when the pattern is common, and the tools in section 4 exist because of this.
Or to the live web, unless the tool around it adds that. Out of the box it knows nothing about your paper.
An invented citation is a sequence of plausible tokens. “Smith and Jones (2019)” is very plausible.
A very well-read, very fast colleague who has read almost everything but remembers none of it precisely, has no access to your files, will never admit to not knowing, and is unapologetic about being wrong.
Context improves AI tools tremendously, and the goal is to provide the right context.
Everything else today is a way of doing that, or of checking what came back.

GPQA Diamond accuracy by model release date, 264 results; each dot is one model. Dashed lines: random guessing (25 percent) and expert human level (about 70 percent). Epoch AI, CC-BY.

Links to Project APE and López-Lira’s system are on the workshop resources page.
Galiani et al. (2026) benchmark applied econometric tasks in Stata, R and Python. Moving from a chatbot that writes one script to an agent that runs and revises its own code is what raises task success.
Under the chatbot, results differ a lot by software. Under the agent, the differences largely disappear.
74 → 96% task success, chatbot to constrained agent
8 cents extra cost per run
Galiani et al. (2026), NBER working paper 35588, abstract. “Constrained agent”: one that runs its code and fixes errors, but within a fixed set of tools.
Two tasks that look equally hard to you can sit on opposite sides of the line.
Drafting prose, summarising a literature, re-deriving a table: inside. But some things are outside, and figuring out which is difficult ex-ante.
Consultants in Dell’Acqua et al.’s field experiment did better inside the frontier and worse outside it, because they trusted the tool in both places.
Dell’Acqua et al. (2026); Ethan Mollick’s “jagged frontier”.

Who has used ChatGPT or Claude inside an IDE or terminal? Who has used Claude Code or the Claude desktop app?
What share of economists in a survey of 1,260 social scientists use a coding agent?
Open dots: anyone using AI at all.
Four in five use some AI tool. One in five uses an agent. Economists are still under forty percent.
Most use chat only
The caveat that this was a really long time ago (February/March).
Lyttelton et al. (2026), survey of 1,260 social scientists, February–March 2026.
About a quarter of PhD students use a coding agent. Fewer than one in ten full professors do.
Different readings: seniors have less to gain, seniors have RAs who use agents, or seniors face the highest switching cost.
Lyttelton et al. (2026).
Code is nearly universal among agent users.
More than half of agent users draft with it, fewer than a third of other users do.
Lyttelton et al. (2026).
Same research, done faster
Faster coding. Stata, R and Python drafts, debugging error logs, robustness checks, replication packages.
Better writing. Structured critique, edits in your own voice, pre-submission referee reports.
New outputs, new projects
New ways to communicate a paper. An interactive website, a policy brief, a dashboard or an audio summary, all from the same paper.
New kinds of projects. Mass automation (Project APE), automated replication and extension (Schwartz et al. 2026), AI agents as simulated humans (Expected Parrot), ???
1 · Research
Mental model: a feedback machine.
2 · Coding
Mental model: a talented RA.
3 · Practical and admin
Mental model: a personal assistant.
Friction is highest. The cost of starting dominates the cost of doing.
The quality bar is lower. No referee, no audience: a funder checks the contents.
There is nobody else to delegate to. Small tasks are not worth giving to someone else.
Different funders, different required formats.
Without an agent. Reformat, re-order, drop sections, rewrite the bio.
With an agent. Drop the master CV and the funder’s template in the folder. The agent reformats, you check.
A task that is easy to verify and that nobody pays you for.
What the expert brings
Better questions. An expert asks the question that is worth asking.
Steering. An expert knows which follow-up moves the conversation where it should go.
Validation. An expert can tell when the output is wrong.
What the tool does not
It does not supply the question. It lowers the cost of asking one seriously.
It does not replace judgment. You decide what to do and whether it is worth doing.
It does not know your institution’s rules unless you put them in the context: the journal’s AI policy, the funder’s format.
Your prompt goes to the model together with the context: CLAUDE.md, a file of standing instructions the agent reads every session, plus the folder listing and the conversation so far.
The model decides which tool to call. The harness runs it and puts the result back into the conversation. Repeat until the task is done, or the agent asks you a question.
Stylised trace of the website session. After Korinek (2025); figure adapted from Velikov (2026). All six tools are in the appendix.
Everything the model conditions on has to fit in one window: the standing instructions, the tool definitions, the conversation, and every file it reads. Drag the slider to the right.
Point at files
@filename
Name the file you mean; do not let the agent read everything in the folder.
One task, one session
/clear
Start a new session for a new task. Thirty turns or a shifted task: restart. Signs of drift.
Watch the meter
/context
Shows what is filling the window and by how much.
Convert once
/pdf-to-markdown
Long PDFs become markdown once, up front; the markdown is committed and pointed at from then on. Which formats need it.
The commands are Claude Code’s; Codex has equivalents. The starter CLAUDE.md template is in the appendix.
The frontier labs
Claude, ChatGPT and Gemini have each led at some point in the last year. Grok (xAI) and Llama (Meta) compete. The leading provider changes with each release.
Open weights
Cheaper to run, and many are not far behind on performance. Mistral is European, which can matter when a data-protection office asks where the model runs.
Switching costs are low. It is more important to start.
Free
$0
Casual chat on claude.ai; no real Claude Code usage
Pro
$20 / month
Light daily Claude Code use, projects, researchStart here
Max
$100–200 / month
Heavy Claude Code use (5× or 20× the Pro limits); the frontier model all day
API
pay per token
Automation, batching, custom-built tooling
Claude’s tiers as of August 2026; OpenAI’s are similar.
Fast, cheap
Haiku
Conversions, summaries, bulk tasks
Default
Sonnet
Daily coding, editing, drafting
Frontier
Opus
Hard reasoning, adversarial review
Most capable, slowest
Fable
Long-running agents, the hardest tasks; costs and waits more
Claude’s model tiers as of August 2026.
Effort sets how many tokens the model spends thinking before it answers. It is a setting in the tool, next to the model choice.
Turn it up for deep analysis and adversarial review. Turn it down for speed on routine tasks.
3–4 characters of English per token
15–30k tokens in a short paper
Tokens are the unit models bill and reason in. Thinking tokens and file tokens come out of the same window.
Rules of thumb from the source deck.
Chat (ChatGPT, Claude.ai)
When to use: brainstorming, polishing one paragraph, quick lookups.
An agent (Claude Code, Codex)
When to use: any task that touches more than one file or needs to execute code.
What chat borrowed from agents
Claude and ChatGPT both have projects: upload files, set standing instructions, keep a memory per project. That is the context half of what an agent has, in a chat window.
Both apps now contain an agent: the Code and Cowork tabs in the Claude app, Codex inside ChatGPT.
What chat still lacks
The loop. Projects hold context; they do not read your folder, write files, or run code and fix what broke.
| surface | what it is good for |
|---|---|
| VS Code | What I use. File tree, diff view, integrated terminal, side-by-side PDF preview. |
| Terminal | claude or codex in any shell. Good for servers and quick one-shots. |
| Desktop app | Claude Desktop’s Cowork and Code tabs, or the Codex desktop app. |
| Web, cloud | claude.ai/code or chatgpt.com/codex. Point at a GitHub repository. |
| Mobile | Capture ideas; monitor or redirect cloud agents. Not for editing. |
| Other editors | Cursor, JetBrains, Zed. Keep your editor, add the agent. |
More at claesbackman.com/agentic-ai-overview.
Both will read your files, run code, edit drafts, iterate, and hand sub-tasks to copies of themselves. A skill is a saved prompt invoked by name.
Standing instructions: CLAUDE.md
Skills trigger: /review-paper
Backed by: Fable, Opus, Sonnet, Haiku
Surfaces: VS Code, CLI, web, desktop
Standing instructions: AGENTS.md
Skills trigger: $review-paper
Backed by: GPT family
Surfaces: VS Code, CLI, web, mobile
Context improves AI tools tremendously, and the goal is to provide the right context.
Chat: you paste it. Projects: you upload it once. An agent: it goes and reads it.
Using the ECB Data Portal API, download two quarterly series from 2000 onwards: the euro-area residential property price index and the deposit facility rate averaged to quarterly. Look up both series keys first and show me the keys and your planned requests before you download anything.
Save the merged series as data/raw/ecb_hp_rate.csv and write data/README.md giving each series key, its units, and its transformation. Use Python.
Then plot the two series together, and report the 2008Q1 and 2022Q1 values of each so I can check them against the portal.
Three paragraphs, three moves: plan first, document, verify. What the prompt is testing for is in the appendix.
And how it is relevant for your own work. Attach the PDF.
From a public API: ECB, FRED, Eurostat, World Bank. Plan first.
For a paper, one step at a time, running each step.
Slides with 2015 examples get current ones; the structure stays.
In your own voice.
Your students may already be doing this.
Agents are fast, cheap and competent at code, including in languages you do not know.
Agents are very good at finding and fixing what crashes. Had an agent used the wrong fixed effect in your table, the code would still run and look plausible, but it would be wrong.
claesbackman.com/verifying-llm-output; Litt (2026); Goldsmith-Pinkham (2026); Scott Cunningham.
The code is the analysis. Once a result is in the paper, you are responsible for it.
Not all code is equal. You don’t need to know the website code, but you should probably check the regressions.
A reviewer that sees everything
Spreads its attention thin, gives generic answers, and spends the window on files that did not change.
A reviewer that sees the diff
Knows exactly what is new and compares it against code you already checked.
You can use git diff, or your agent can. A diff is the list of lines that changed since the last saved version.
Version control is a verification tool. This is a good reason to use Git.
Two ways to get a reviewer with no stake in the code:
Ask the current agent to hand the diff to a subagent, a fresh copy of itself with an empty context.
Open a new session and point it at the diff.
Look at what the agent found and verify whether you agree with them. .
Weakness: a fresh copy of the same model shares its habits. Hence check 2.
Two copies of one model share training data and habits of thought, so they make the same mistakes. Send the same diff to a model from a different developer.
Open a fresh session and ask:
Read the last commit and the files it touches, then ask me three questions about what it changed. Do not tell me the answers until I have answered. If my answer is wrong, say so directly; do not tell me I was close.
Always a new session, not the one that made the change. Starting fresh forces the agent to read the files instead of recalling the conversation.
A skill for it: /explain-diff writes one offline page about what changed, why, what it does to the results, and ends with a five-question quiz. More skills in the appendix.
What and why
If your code is Stata, a second agent writes the results in Python or R from the paper’s description alone. An idea from Scott Cunningham.
If coding slips are independent across languages, differences in the output reveal the errors.
How to isolate it
New folder, fresh session, a spec.md copied from the paper with the results deleted, a README.md describing the raw data.
This is a major enterprise, so reserve it for the main tables and for things that matter.
Two dot per student, each is an exam.
Almost everyone scored above 90 on the take-home midterm.
On the in-person final the same students spread from 95 down to zero.


+18% homework score for pupils using AI
−30% homework completion time
Distributions for pupils in China: never used AI, before using AI, and using AI. Charts reproduced as published; study as cited in Bryan (2026).

−20% exam performance for the same pupils
Homework up, time down, exams down. The homework stopped teaching.
Same study and source as the previous slide.
40 → 27 hours a week of study, US college students, 1961–2003
one in four finance students copied wrong Chegg answers, and learned less
Both figures as cited in Bryan (2026).
Bryan (2026), kevinbryanecon.com/teachwithai.
Rules 1–3 of Bryan (2026).
For the students
Help them learn more efficiently. Class-specific tutoring, spaced repetition, mastery learning at scale.
Personalise assignments. Problems at the level each student is struggling with.
For the teacher
Improve your own teaching. Find out where you were confusing before the exam tells you.
Raise standards. If research and drafting are cheaper, the bar moves up.
Rules 4–7 of Bryan (2026).
A shared assistant (a Claude project, custom GPT or NotebookLM) loaded with your lecture notes and problem sets.
Standing instructions: use the course’s notation, never give the answer, ask the student to explain first, return to topics they missed.
Each week, export the questions students asked, anonymised.
Prompt: “Group these by underlying misconception, rank by frequency, and point to the lecture slide most likely to have caused each.”
One problem, three difficulty tiers. Assign the tier by last week’s quiz result.
Prompt: “Write this problem at three levels: scaffolded, standard, extension. Same concept, full solutions, and a rubric.”
Student questions are personal data. Anonymise before exporting and check the institutional policy once.
The researchers who get the most out of these tools treat the AI as a colleague to argue with, which is what the mental model in section 1 said.
| habit | what to look for | how |
|---|---|---|
| Run the code | Plausible Stata, R or Python that does not always run | Run it yourself |
| Check the numbers | Statistics, dates and quotes get invented | Re-derive from the table or the paper |
| Read the citations | Citations to papers that do not exist are common | Look for it on Google Scholar |
| Re-read the passage | Confident descriptions of text that is not in the draft | Open the file, read the section |
Adapted from Korinek (2025).
Disclosure
Bryan’s rule 8: write the AI policy into the syllabus, per assignment.
Data handling
Check once, before the first paste; Day 2 has the secure-server workflow.
Publishing gets noisier
Journals fill with polished, competent-looking mediocrity. Signal-to-noise drops and reviewing gets harder. Editors and referees absorb the cost.
Research can still get better
Ideas that could not be written before reach readers. The language barrier falls. Feedback that used to need a good department is available to anyone with a folder and twenty dollars.
The open question: how do we evaluate research once competent-looking writing is free?
Later we will use agents together.
Hands-on with your own materials: a referee report on your draft, a voice file from your own writing, your code and your paper side by side.
Bring a draft you care about and four or five PDFs of your own writing. Slides, skills, guides: claesbackman.com/ai_workshop
Claes Bäckman, SAFE
| from | slide | what is in it |
|---|---|---|
| 4 · Harness | CLAUDE.md starter template | Standing instructions to adapt for your own project |
| 4 · Harness | The six tools the harness adds | Read, edit, shell, search, web, MCP |
| 4 · Harness | Signs of session drift | When to start a fresh session |
| 4 · Harness | Which file formats need converting | Text, convert once, read through code, OCR |
| 4 · Harness | Converting is one sentence | A prompt that converts a folder of PDFs |
| 6 · In practice | What the ECB prompt is testing for | APIs, series keys, and what to check |
| 7 · Verification | My skill library | Nine skills on GitHub, free to fork |
| 7 · Verification | Four checks, the full table | What each check catches and what it costs |
| References | Everything cited on the slides |
# About me
I am a researcher in [field].
# How I want you to work with me
- Ask clarifying questions before generating long output.
- Critical, skeptical tone in feedback. Do not flatter.
- When editing, preserve my voice (see voice.md).
No generic LLM phrasing.
- Avoid bullet lists and passive voice in formal writing.
- Cite the specific line or section you are commenting on.
# Things to avoid
- Do not fabricate citations.
- Do not over-claim causality.
- Do not insert emoji or markdown decorations in formal docs.Always loaded, every session. Standing instructions, the things you would otherwise repeat in every prompt.
Day 2 builds one for your own project with /init.
Back to the harness, the habits or the appendix overview.
| capability | what it lets the agent do |
|---|---|
Project files (Read) |
Read files, inspect folders, understand the project structure |
Editing (Edit, Write) |
Propose and apply changes; create or modify files |
Shell (Bash) |
Run commands such as Rscript, pdflatex, tests, linters |
Search (Grep, Glob) |
Search file names and file contents in the project |
Web (WebSearch) |
Search or fetch web pages when web access is enabled |
| MCP and plugins | Connect to external tools (GitHub, databases, custom servers) through a standard interface, the model context protocol |
The names are Claude Code’s; Codex has the same six under other names. Back to the loop or the appendix overview.
CLAUDE.mdFix. Start a fresh session or close and reopen. Re-state the task. Re-attach the relevant file.
Thirty turns, or the task has shifted: start fresh. Always.
Back to the habits or the appendix overview.
Native, no friction
.md, .tex, .txt, .bib, .csv, .json, .yaml, and code: .py, .R, .do, .ipynb
Convert once
.docx, born-digital .pdf, .xlsx: to markdown or CSV, saved next to the original
Read through code, not as text
Data binaries such as .dta and .rds: the agent loads them in Stata, R or pandas and looks at the output
Hard: OCR or skip
Scanned PDFs, survey exports such as .qsf
Convert once, commit the markdown alongside the original, point the agent at the markdown.
Back to the habits, on to converting, or the appendix overview.
“Convert every PDF in papers/ to markdown, save each next to the original, and tell me which ones came out garbled.”
Back to the file formats or the appendix overview.
RPP.Q.I9.N.TD.00.3.00. The agent must look the key up before it can fetch anything, which is why the prompt asks to see the keys first.Back to the prompt or the appendix overview.
| skill | use |
|---|---|
/review-paper |
Pre-submission referee report, 8 agents |
/review-paper-light |
Fast 2-agent version, about 1 minute |
/review-paper-code |
Paper and code alignment review |
/review-grant |
Pre-submission grant review, 6 agents |
/review-pap |
Pre-analysis-plan review |
/audit-analysis |
Adversarial audit of changed analysis code |
/explain-diff |
Explain a code change as an HTML page with a quiz |
/paper-version |
Paper to policy brief, 1-page or 5-page summary |
/pdf-to-markdown |
Convert PDFs to readable markdown |
All at github.com/claesbackman/AI-research-feedback (MIT), linked from claesbackman.com/ai_workshop. Fork them, modify them, make them yours. Back to check 3 or the appendix overview.
| check | what you do | what it catches | cost |
|---|---|---|---|
| Fresh agent attacks the diff | A fresh chat or subagent reviews the diff | Errors invisible to the agent that wrote the code | Minutes |
| Different model | Same diff, model from another developer (Codex vs. Claude) | Shared habits: two copies of one model make the same mistakes | Minutes |
| Make the agent quiz you | Fresh session explains the change, then quizzes you | Your own understanding gap | 10–15 min |
| Reimplement in another language | Fresh folder, spec.md and README.md only, other language and model, you run it |
Code that faithfully implements something other than the paper | Hours |
Back to the four checks or the appendix overview.