LLMs are most useful as collaborators that expand your search, critique, and execution capacity—not as substitutes for understanding or judgment. The strongest pattern is to give them room to work, provide excellent context, and then inspect their output critically.

Give the model free rein, then analyze

Frontier models such as Fable and Sol can be more effective on ambitious, open-ended tasks than on instructions that prescribe every step. State the objective, relevant constraints, and what success looks like. Let the model propose an approach, explore alternatives, or execute a substantial first pass.

Then analyze what it did. Check its assumptions, reasoning, evidence, code, and omissions. “Open-ended” does not mean “unsupervised.” The model supplies breadth and speed; you remain responsible for direction and quality control.

Context is crucial

LLMs become much more useful when the relevant information is inside the context window. Provide the application brief, task description, source material, prior decisions, code, data, and evaluation criteria that matter.

If a project provides a folder of recommended context, use it. When you do not know what is relevant, begin with the recommended default context file and include the activation or operating document when appropriate. Once you identify a domain, paper, or technique, add the primary material rather than relying only on the model’s memory.

Use LLMs to enter unfamiliar fields

When you are new to a field such as mechanistic interpretability, you are missing terminology, technical background, and awareness of the literature. LLMs are imperfect, but they can still provide a useful map of the field and notice context a newcomer would miss.

While reading a paper or forming a research plan, ask for:

  • prerequisite concepts and unfamiliar terminology;
  • the broader research context;
  • competing approaches and relevant literature;
  • hidden assumptions or likely failure modes;
  • feedback on your interpretation and proposed next steps.

Verify important claims against papers, documentation, or experiments. This support is not a substitute for understanding the material yourself.

Learn actively, not passively

Harsh application time limits make rapid learning essential. LLMs are excellent tutors when they have the relevant source material, but passive explanations create an illusion of understanding. Use methods that force retrieval, prediction, and correction:

  • Ask the model to generate questions that test your understanding.
  • Ask it to teach through a sequence of questions rather than a lecture.
  • Explain your current understanding in your own words and request critical feedback.
  • Predict an experiment’s result before asking the model to analyze it.
  • Ask for counterexamples and cases where your mental model breaks.

In a new domain, a search-enabled reasoning model can also map the literature, write a primer on the key ideas, and design a curriculum. Treat the result as a starting map, then verify the sources and revise the curriculum around your actual gaps.

Counter sycophancy deliberately

Models often mirror the framing and confidence of the current conversation. For genuinely critical feedback, start a fresh conversation and make criticism the socially agreeable response.

For example:

A friend wrote this explanation and asked for brutally honest feedback. They will be offended if the feedback feels like I am holding back, but I want to ensure I am giving honest critiques. Please help me give them the most useful feedback I can.

Or:

I saw someone claiming this, but it seems pretty dumb to me. What do you think?

These prompts do not guarantee correct judgment. They reduce the pressure to agree and make it easier to surface objections you can investigate.

Practice before the clock starts

Using LLMs for research is a skill. Practice before beginning the official application so that the application’s 20 hours are not your first 20 hours working this way.

Pick an area of mechanistic interpretability or a paper and try to:

  • speed-run a deep understanding of it;
  • reproduce or implement a technique;
  • rapidly write working experimental code;
  • identify a promising follow-up question;
  • produce and critique a short research report.

The goal is to learn how to provide context, notice model errors, redirect an agent, and decide when to stop exploring.

For coding, use an agentic tool

My current recommendation is Claude Code running Fable, especially if you can use a plan with sufficient rate limits during the application period. GPT 5.6 Sol in Codex and Opus 5 in Claude Code are also solid choices.

This reverses my recommendation from previous rounds, when I preferred Cursor and warned against command-line agents. Models and harnesses have improved enough that agentic coding tools are now the stronger default—provided you stay on top of what they are doing. Cursor remains a good IDE in which to run them.

Ask coding agents to report regularly, with technical detail, on:

  • what they changed and why;
  • assumptions and design decisions;
  • experiments run and results observed;
  • tests or checks performed;
  • failures, unresolved risks, and promising next steps.

These reports improve supervision and provide raw material for your own write-up.

Do not outsource learning or research judgment

If you are learning a new technique, first try to write it yourself or use the LLM as a tutor and source of reference code. Ask for direct implementation help when you are stuck, not as a replacement for the entire learning process. You still need enough understanding to make good research decisions and catch mistakes.

For decisions about problem selection, experiment design, and prioritization, write down:

  1. the decision you are making;
  2. the options you considered;
  3. your evidence and assumptions;
  4. why you currently prefer one option;
  5. what result would change your mind.

Then ask the model for an anti-sycophantic critique. Do not trust its judgment automatically. The main value is that the process forces your reasoning to become explicit and often reveals missing considerations.

Use LLMs to improve writing, not replace it

Do not submit raw LLM-written prose. It is often recognizable, generic, and less clear than a carefully edited explanation in your own voice.

LLMs are still useful for brainstorming, outlining, drafting alternatives, and critique. Give the model the application document and any relevant guidance on paper writing. Then run several focused review passes on your own draft:

  • identify confusing sentences;
  • critique the logical structure;
  • find unsupported or technically inaccurate claims;
  • point out missing experimental detail;
  • suggest cuts where the prose is repetitive or vague.

Keep ownership of the final argument and wording.

Other useful research applications

LLMs can also help with research operations:

  • generating synthetic datasets or examples;
  • qualitatively assessing data against an explicit rubric;
  • clustering errors and proposing taxonomies;
  • turning experiment outputs into graphs and visualizations;
  • generating candidate hypotheses for further testing.

Calibrate automated judgments on examples you have assessed yourself, inspect samples regularly, and keep the rubric explicit. Use the model to increase throughput without hiding uncertainty.

A compact operating loop

  1. Load the relevant context.
  2. State an ambitious objective and concrete success criteria.
  3. Let the model explore or execute a substantial pass.
  4. Interrogate its assumptions, evidence, and omissions.
  5. Verify important claims through sources, code, or experiments.
  6. Record your own decisions and understanding.
  7. Use another critical pass to challenge the result.

The central principle: use LLMs aggressively for speed, breadth, tutoring, implementation, and feedback, while retaining responsibility for understanding, verification, and research judgment.