Supervising Students
I'm generally excited to supervise any student projects at Cambridge for the Part II, Part III, or MPhil.
If you would like to discuss a potential project, please get in touch.
Project Ideas
I've collected a set of potential student projects ideas below:
- J-Space Manipulation and Attack Detection - Recent research has identified a small collection of internal neural patterns in large language models, called the J-space, that appears to support silent reasoning and reveal concepts a model is considering without expressing them in its output. The purpose of this project would be to investigate attacks against this mechanism. In particular, the project would search for categories of model input that disproportionately disrupt the J-space, suppress or redirect its representations, or otherwise interfere with its role in reasoning. Candidate inputs could include existing attacks such as imperceptible perturbations and cross-prompt injection attacks (XPIAs), as well as new adversarial inputs developed during the project. A successful project would produce a systematic evaluation of how different attacks affect J-space signals and downstream model behavior, identify whether J-space disruption predicts attack success, and potentially develop new attacks or J-space-based detection and defense techniques. Because this project builds on early research, it is not clear in advance which approaches will succeed; strong results could nevertheless have significant implications for LLM security and interpretability. This project would likely be a good fit for Part III / MPhil students, or in partial form as a Part II project.
- AI Watermarking Attacks Claimed - From 2 August 2026, Article 50(2) of the EU AI Act requires providers of AI systems that generate synthetic text to ensure that their outputs are marked in a machine-readable format and detectable as artificially generated or manipulated. Although this requirement applies to providers placing systems on the EU market and to certain uses of their outputs within the EU, its effects may extend worldwide. For example, Claude's watermarking will be applied globally at launch. Claude uses a version of Google DeepMind's SynthID-Text technique: when several next words are similarly appropriate, a secret key and the preceding words determine the source of randomness used to choose among them, creating a statistical pattern that can later be detected with the key. The general technique and a reference implementation are public, although Claude's key and exact production configuration are not. The purpose of this project would be to investigate attacks against this watermarking mechanism. Possible directions include developing methods to remove text watermarks while preserving meaning and readability, forging watermarks in human-written text to create false attributions, and exploring other forms of spoofing or manipulation through paraphrasing, backtranslation, or text mixing. A successful project would produce an automated attack and evaluation tool and measure how reliably watermarks can be removed or forged without noticeably changing the text. This project would likely be a good fit for Part III / MPhil students, or in partial form as a Part II project.
- LLM Vulnerability Detection - Traditionally automatic vulnerability detection in code occurs via static code analysis and fuzzing. These techniques tend to be very good at detecting specific classes of technical vulnerabilties such as use-after-free bugs. However, these techniques are less adept at detecting logical application errors, such as detecting whether the correct AuthN model has been used. The advent of large language models (LLMs) may offer a better way to detect these attacks. The purpose of this project would be to investigate whether LLMs can be used to identify logical vulnerabilities in application software implementations. A successful project would produce a literature review, a theoretical analysis of the problem space, and experiments evaluating LLMs' abilities to detect logical vulnerabilities. Excellent projects may choose to fine-tune models as part of experimentation and submit results to relevant security/ML conferences. This project would likely be a good fit for Part III / MPhil students.