Computer Laboratory

MPhil, Part III, and Part II Project Suggestions (2026-2027)

Contact

 

 

 

 

 

 

 

  

 

Please contact Eiko Yoneki (email: eiko.yoneki@cl.cam.ac.uk) if you are interested in any project below.

 

1.       Evaluating Frontier Models for AI Infrastructure +++

 

Contact: Eiko Yoneki

 

This project will explore InfraBench, developing benchmarks to evaluate frontier the ability of AI models to design, implement, optimise, and evolve AI infrastructure. The scope will span multiple layers of the AI systems stack, including kernel optimisation, inference serving, training systems and algorithms, and broader software-engineering tasks involving complete end-to-end systems.

 

A particular focus will be on self-evolving AI infrastructure: investigating whether LLM-based agents can autonomously identify performance bottlenecks, propose and implement system modifications, validate functional correctness, and iteratively improve performance using execution traces, profiling data, and runtime feedback. Rather than evaluating one-shot code generation alone, the project will study models operating within a closed-loop optimisation process involving generation, execution, measurement, diagnosis, and refinement.

 

The project will develop a diverse set of realistic tasks and reproducible evaluation methodologies that measure functional correctness, computational efficiency, optimisation effectiveness, and the quality and robustness of system-level improvements. The longer-term goal is to understand how effectively frontier models can move beyond code generation towards autonomously engineering and optimising complex AI systems.

 

Desirable skills:

-          Basic understanding of MLSys and AI infrastructure, including LLM training, inference, and associated algorithms.

-          Experience in, or interest in, building efficient infrastructure for LLMs or other workloads and hardware.

-          Familiarity with hardware stacks, GPU programming, or performance profiling would be beneficial.

 

References:

[1]    Ding, L. et al. (2026). Φ-Bench: Can Large Language Models Engineer the Infrastructure That Powers Them? arXiv:2609.10226.

[2]    Xing, S. et al. (2026). FlashInfer-Bench: Building the Virtuous Cycle for AI-driven LLM Systems. MLSys 2026.

[3]    Zhang, G. et al. (2026). AccelOpt: A Self-Improving LLM Agentic System for AI Accelerator Kernel Optimization. MLSys 2026.

[4]    He, X. et al. (2026). SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories? ICML 2026.

 

 

2.       Optimising KV Cache Management for Coding Agents +++

 

Contact: Eiko Yoneki

 

This project will investigate new abstractions and algorithms for managing KV caches in coding-agent workloads. Unlike conventional LLM serving, coding agents involve long-running, iterative, and potentially branching executions, with frequent pauses for tool calls and subsequent resumptions. Their KV states may reside in GPU HBM, CPU DRAM, or SSD, or may be discarded and recomputed when needed. Consequently, request admission, execution scheduling, and KV-cache placement and movement become tightly coupled system-level decisions.

 

The project will explore how these decisions can be coordinated across multi-step agent executions, taking into account memory capacity, data-transfer bandwidth, execution dependencies, and expected future reuse. Potential directions include policies for KV reuse, offloading, eviction, prefetching, migration, and recomputation, as well as scheduling mechanisms that exploit idle periods during tool execution.

 

The goal is to develop a unified approach to request scheduling and KV-cache management that exploits the structure and lifecycle of agentic workloads to reduce latency, increase throughput, and improve overall resource efficiency.

 

Desirable skills:

  • Basic understanding of MLSys and AI infrastructure, including LLM inference, attention, KV caching, and associated algorithms.
  • Experience in, or interest in, building efficient infrastructure for LLMs or other workloads and hardware.
  • Familiarity with hardware stacks, GPU and CPU memory hierarchies, storage systems, or performance profiling would be beneficial.

 

 

References:

[1]     Liu, Y. et al. (2025). LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference. arXiv:2510.09665.

[2]     Chen, Q. et al. (2026). CONCUR: High-Throughput Agentic Batch Inference of LLM via Congestion-Based Concurrency Control. ICML 2026.

[3]     Li, H. et al. (2025; revised September 2026). Continuum: Efficient and Robust Multi-Turn LLM Agent Scheduling with KV Cache Time-to-Live. arXiv:2511.02230.

[4]     Bian, Z. et al. (2025; revised September 2026). TokenCake: A KV-Cache-centric Serving Framework for LLM-based Multi-Agent Applications. arXiv:2510.18586; accepted at EuroSys 2027.

[5]     Qiu, S. et al. (2026). Tutti: Making SSD-Backed KV Cache Practical for Long-Context LLM Serving. arXiv:2605.03375.

 

 

3.       Learning-Guided GPU Kernel Optimisation: RL and LLM Code Generation +++

 

Contact: Eiko Yoneki

 

Creating high-performance GPU kernels is challenging. Closed-source GPU compiler backends make it difficult to understand and tune code generation for specific hardware, while compiler-generated assembly schedules are not always optimal for performance-critical kernels. This project investigates two complementary learning-guided approaches to GPU optimisation: low-level reinforcement-learning (RL) based assembly scheduling and LLM/agent-based kernel generation.

The first approach formulates assembly optimisation as a search problem. CuAsmRL [1], developed by our group, uses deep RL and runtime feedback to optimise NVIDIA CUDA SASS instruction schedules through binary rewriting. It replaces labor-intensive manual scheduling with automated exploration of hardware-specific schedules.

The second approach uses LLMs and coding agents to generate and iteratively optimise GPU kernels. Recent systems typically operate at CUDA, CuTe, or Triton level and use compilation, correctness testing, profiling, and runtime feedback. LLMs can propose higher-level semantic and algorithmic transformations, whereas RL/search can explore a narrower hardware-specific scheduling and configuration space. Both approaches are based on learning-guided optimisation as base.

The project explores the gap between two approaches by benchmarking. LLMs provide semantic transformations, while RL/search can explore the much narrower hardware scheduling/configuration space.

You can start understanding and replicate CuAsmRL. Then study existing LLM based code generation tools to select appropriate one for the comparison. Possible candidates are as follows. You could choose more than one if time allows.

1.       KernelPro - Multi-agent LLM optimiser generating CUDA/CuTe while using SASS binary analysis to diagnose instruction-level bottlenecks; it does not directly rewrite SASS. arXiv

2.       Harness Engineering / Codex-Claude kernel agents — Uses Codex and Claude Code with profiling feedback to generate optimised Blackwell kernels, but optimisation occurs at kernel source level rather than direct SASS transformation. arXiv

3.       SASS-Bench - Benchmark for evaluating model/agent understanding of actual NVIDIA SASS execution, potentially providing infrastructure for future SASS-generating agents, but is not itself an optimiser. GitHub

4.       sasskit / Reforge - Directly performs empirical SASS→SASS post-compilation optimisation on Blackwell binaries, but currently uses mutation/search rather than an LLM agent. GitHub

 

The longer-term goal is to understand whether these approaches can form a hybrid learning-guided compilation pipeline. If time permits, the project can extend CuAsmRL towards an LLM/agent-driven SASS optimiser in which an agent inspects SASS and profiling information, proposes low-level transformations or scheduling decisions, verifies correctness, benchmarks the resulting binary, and iteratively improves it.

Desirable skills: CUDA, RL, GPU, LLM

 

References:

[1] G. He and E. Yoneki: CuAsmRL: Optimizing GPU SASS Schedules via Deep Reinforcement Learning. 2025.  https://www.cl.cam.ac.uk/~ey204/pubs/2025_CGO.pdf

 

4.       Agent-Guided Kernel Optimisation for Tenstorrent Accelerators +++

 

Contact: Eiko Yoneki

 

Generating high-performance accelerator kernels is difficult because performance depends on low-level architectural decisions such as tiling, memory placement, data movement, pipelining, and parallel mapping. Recent work such as CAKE [1] suggests that LLM-based kernel optimisation can be improved by exposing compiler and hardware knowledge to the optimisation agent rather than treating compilation and execution as a black box. However, this approach has primarily been explored for GPU architectures, and its applicability to spatial/dataflow-oriented accelerators remains largely unexplored.

This project will investigate agent-guided kernel optimisation for Tenstorrent accelerators using the TT-Metalium programming model [2,3,4]. Tenstorrent explicitly exposes reader/compute/writer kernels, circular buffers, local L1 memory, NoC communication, and multi-core mapping, providing a useful platform for studying architecture-aware optimisation. The student will build an optimisation loop in which an LLM agent modifies TT-Metalium kernels and receives correctness and performance feedback. A baseline using only compilation, correctness, and execution-time feedback will be compared against a proposed approach that additionally provides structured Tenstorrent-aware information, including circular-buffer/resource usage, data movement, core mapping, runtime failures, and profiler-derived bottlenecks.

The main research question is whether structured hardware/compiler feedback enables the agent to find correct and high-performance implementations more efficiently. Evaluation will use a small portfolio of TT-Metalium kernels, including elementwise operations and matrix multiplication, and measure final speedup, valid-candidate rate, optimisation cost, convergence, and generalisation across input shapes. Ablation studies will identify which feedback signals are most useful. If recurring optimisation decisions emerge, a stretch goal is to derive a small Tenstorrent-specific schedule/action representation. The broader objective is to examine whether compiler-agent co-design can generalise beyond conventional GPUs to spatial/dataflow accelerator architectures.

If the project succeeds, a follow-on work could add NVIDIA GPU + Tenstorrent to test whether the same agent/compiler-feedback methodology transfers across conventional GPU and spatial/dataflow architectures.

Desirable skills: kernel Optimisation, Tenstorrent

 

References:

[1] Ye et al., CAKE: Compiler-Agent Co-Design for Frontier Kernel Evolution, arXiv:2608.12629, 2026.

[2] Tenstorrent, TT-Metalium Getting Started, TT-Metalium Documentation, 2026.

[3] Tenstorrent, TT-Metalium Programming Examples, TT-Metalium Documentation, 2026.

[4] Tenstorrent, Memory from a Kernel Developers Perspective, TT-Metalium Documentation, 2026.

 

5.      Verification-Guided LLM Optimisation: Combining LLM-as-a-Judge with Equivalence Checking++

 

Contact: Eiko Yoneki

 

Large language models are increasingly used to generate and optimise programs, compiler transformations, and accelerator kernels. An LLM-driven optimisation loop can propose a faster implementation, benchmark it, and repeatedly refine the result. A central challenge, however, is correctness: an optimisation is useful only if the transformed program preserves the behaviour of the original.

 

LLM-as-a-Judge provides a lightweight way to assess whether two implementations appear semantically equivalent, but such a judgement is probabilistic rather than a proof. In contrast, testing and equivalence-checking tools such as Alive2 and VOLTA can provide stronger evidence or formal guarantees for supported program classes. This project investigates how these mechanisms can be combined inside an iterative LLM optimisation loop.

 

The project will design and evaluate a small verification-guided optimisation framework. An LLM first proposes an optimisation. An LLM judge then acts as a cheap semantic filter before the candidate is subjected to testing and, where possible, formal equivalence checking. Correctness and performance feedback are returned to the LLM to guide subsequent optimisation attempts.

 

1)      Original program → LLM optimiser → Candidate

2)      LLM-as-a-Judge → Testing → Equivalence checking → Benchmarking

3)      Feedback to optimisation loop

 

The expected outcome will be a prototype verification-guided LLM optimisation pipeline and an empirical comparison of LLM judgement, testing and formal equivalence checking. The project should clarify where an LLM judge is useful as an inexpensive first-stage validator, where it fails, and whether layered verification can make automated LLM-driven program optimisation more reliable and efficient.

 

Desirable skills: LLM, Optimisation

 

References:

[1]     Dubey et al. Equivalence Checking of ML GPU Kernels. arXiv:2511.12638, 2025. (VOLTA).

[2]     Lopes et al. Alive2: Bounded Translation Validation for LLVM. PLDI, 2021.

[3]     Zheng et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS Datasets and Benchmarks, 2023.

[4]     Alive2 project: https://github.com/AliveToolkit/alive2

 

6.       Probabilistic Inference for Reinforcement Learning +++

 

Contact: Eiko Yoneki

 

Reinforcement learning (RL) relies on hand-crafted reward functions, which are often difficult to design and may not fully capture task goals. Probabilistic inference provides an alternative view, treating decision making as an inference task where rewards can be derived from probabilistic models rather than explicitly specified.

This project will investigate how simple probabilistic models - such as Hidden Markov Models (HMMs) and Bayesian Neural Networks (BNNs) can be adapted for RL tasks. The focus will be on using inference techniques (e.g., expectation-maximisation, variational inference) to derive approximate reward signals from observed trajectories.

You will use existing probabilistic programming and RL frameworks, such as Pyro/PyTorch, NumPyro, or TensorFlow Probability. Experiments will start with GridWorld and may extend to CartPole or MountainCar, comparing the proposed approach with standard RL (e.g., Q-learning, DQN, PPO) and, where appropriate, IRL baselines. Evaluation will consider performance, sample efficiency, reward/objective recovery, robustness to uncertainty, and computational cost. If time permits, the approach may also be evaluated on a computer-systems optimisation case study.

 

Desirable skills: HMM, RL, Probabilistic Model

 

References:

[1] Levine, S. (2018). Reinforcement Learning and Control as Probabilistic Inference: Tutorial and Review. arXiv:1805.00909.

[2] Toussaint, M., & Storkey, A. (2006). Probabilistic Inference for Solving Discrete and Continuous State Markov Decision Processes. ICML.

[3] Koller, D., Friedman, N. (2009). Probabilistic Graphical Models: Principles and Techniques. MIT Press.

 

7.       Bandwidth-Aware Parallelism Planning for Distributed LLM Training and Inference++

 

Contact: Eiko Yoneki

 

Given the various communication bandwidths in cloud environments, for instance, NVLink (800/600/400 GBps, depending on generation), InfiniBand (400 GBps), RoCE (40 GBps), TCP (1-5 GBps), PCIe (20-80 GBps), and Ethernet (<1 GBps) and multiple parallelism strategy choices including data parallelism [1], tensor parallelism [2,3], pipeline parallelism [4,5], expert parallelism, and sequence parallelism [6], it is essential to design an automatic-parallelism method that takes the physical network topology as input and outputs a parallelism plan that minimises communication volume on low-bandwidth communication links and maximises communication volume on high-bandwidth links [7,8,9]. The system should build a network topology-aware optimisation engine that automatically profiles the physical infrastructure and searches through different parameter slice allocation and corresponding communication operations among different slices to minimise overall communication overhead. The key metrics to measure include end-to-end training time, communication-to-computation ratio, network utilisation efficiency across different bandwidth tiers, and performance comparison against existing frameworks like Megatron-LM [10,11].

First, you need to understand the detailed characteristics of different parallelism, then build appropriate modelling, then integrate some of the characteristics into existing simulator, and get simulation results that reflect real-world experimental results, and finally finish the automatic search that could guide real-world deployment.

 

Desirable skills: Distributed systems programming, performance modelling and profiling, network topology analysis

 

References: 

[1] Li, M., et al. "Scaling Distributed Machine Learning with the Parameter Server." OSDI 2014 

[2] Shoeybi, M., et al. "Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism." arXiv:1909.08053, 2019 

[3] Narayanan, D., et al. "Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM." SC 2021 

[4] Huang, Y., et al. "GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism." NeurIPS 2019
[5] Fan, S., et al. "Dapple: A Pipelined Data Parallel Approach for Training Large Models." PPoPP 2021 

[6] Korthikanti, V., et al. "Reducing Activation Recomputation in Large Transformer Models." MLSys 2023 

[7] Zhang, S., et al. "Auto-Parallelizing Large Models with Rhino: A Systematic Approach on Production AI Platform." arXiv 2023

 

 

8.       Multi Objective Scheduling of Distributed LLM Serving on Heterogeneous GPUs ++

 

Contact: Eiko Yoneki

 

Heterogeneous GPU clusters offer a potential solution to mitigate significant operational cost of Large Language Model (LLM) inference. However, existing scheduling systems do not adequately address the unique computational patterns of new Mixture-of-Expert (MoE) models within these complex environments.

 

This project addresses this gap by developing and evaluating a novel scheduling algorithm, specifically designed to optimise MoE model inference on heterogeneous

GPU clusters. The initial work has been explored in [1][2][3], where the algorithm is separated into two distinct processing stages: an outer loop and an inner loop. The outer loop uses Bayesian Optimisation to efficiently search the complex configuration space, partitioning an inventory of different GPUs into small optimal-islands . For the inner loop, this work implements a new linear programming formulation that precisely maps workload ranges, categorised by input sequence length and separated into prefill and decode phases, to the generated islands.

 

In this project, multi objective optimisation (e.g. Multi Objective BO) will be explored such as the latency, accuracy, and/or power consumption. The scheduling algorithm was evaluated using a simulation framework on real-world workload traces.

 

This work will demonstrate a workload-aware scheduling approach can unlock substantial performance and cost-efficiency gains for serving large-scale MoE models in complex, heterogeneous environments.

 

Desirable skills: BO, LLM, GPU, Multi-Objective

 

References:

[1] Nathan Rignall: Scheduling of Distributed LLM Serving on Heterogeneous GPUs, 2025 (https://www.cl.cam.ac.uk/~ey204/pubs/MPHIL_P3/2025_Nathan.pdf).

[2] Y. Jiang et al.: Demystifying Cost-Efficiency in LLM Serving over Heterogeneous GPUs. ICML 2025 (arXiv reprint).

[3] BOute: Cost-Efficient LLM Serving with Heterogeneous LLMs and GPUs via Multi-Objective Bayesian Optimization. MLSys 2026 (arXiv reprint).

[4] S.Alabed:  BoGraph: Structured Bayesian Optimization From Logs for Systems with High-dimensional Parameter Space. 2022. (https://arxiv.org/pdf/2112.08774).

 

 

9.       BO+RL for Multi-Model LLM Serving ++

 

Contact: Eiko Yoneki

 

This project proposes a two-timescale, learning-augmented framework for multi-model LLM serving on heterogeneous GPU clouds. The outer loop performs static optimisation-selecting model placements and replica counts-via cost-aware Bayesian Optimisation that accounts for resource limits, reconfiguration overheads, and workload uncertainty. The inner loop delivers dynamic optimisation with safe reinforcement learning and bandits for online routing, batching, and draft-verify orchestration, informed by output-length prediction and short-horizon demand forecasts. Objectives are explicitly multi-objective: maximise goodput while satisfying latency SLOs (TTFT, end-to-end), and jointly reduce cloud cost, energy, and unfairness using constraints and CVaR/tail controls. The loops are coupled by contracts (budgets, SLO targets) and telemetry (utilisation, tail latency), enabling robust adaptation. Evaluation on real traces and mixed GPU clusters will produce Pareto fronts and ablations (heuristics, BO-only, RL-only, BO+RL), demonstrating scalable, adaptive, and economical LLM serving.

Desirable skills: Python, LLM, BO, RL

 

References:

[1]    Kwon, W. et al. Efficient Memory Management for Large Language Model Serving with PagedAttention (vLLM). SOSP23.

[2]    Zhong, Y. et al. DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized LLM Serving. OSDI24.

[3]    Jiang, Y. et al. ThunderServe: High-Performance and Cost-Efficient LLM Serving in Cloud Environments. arXiv 2025.

[4]    Duan, J. et al. MuxServe: Flexible Spatial-Temporal Multiplexing for Multiple LLM Serving. arXiv 2024.

[5]    Leviathan, Y., Kalman, M., Matias, Y. Fast Inference from Transformers via Speculative Decoding. ICML 2023.

 

 

10.   Multi-Objective Compiler Optimisation ++

 

Contact: Eiko Yoneki

 

[1] demonstrated the feasibility of using Reinforcement Learning (RL) to optimise LLVM pass lists, improving runtime performance over heuristic-based baselines. A natural extension is to move beyond single-objective optimisation and treat compiler design as a multi-objective problem. In practice, developers must balance competing goals such as runtime speed, binary size, and energy efficiency. These objectives often conflict passes that accelerate execution may increase code size or energy usage.

This project will extend [1] s RL framework by defining multi-objective reward functions, enabling agents to explore trade-offs rather than a single performance metric. Approaches include scalarisation (weighted combinations of objectives) and Pareto-based RL, which approximates the set of non-dominated optimisation strategies. Benchmarks from PolyBench, CoreMark, and MiBench will be used to evaluate outcomes against LLVM defaults (-O3, -Os). The result will be a prototype system showing how learning-based compilers can adapt policies across multiple optimisation criteria, aligning with real-world deployment needs.

Desirable skills: RL, Compiler

 

References:

[1] Yilin Sun: Optimizing LLVM Pass List using Reinforcement Learning.

[2] Mammadli: A. Reinforcement Learning for Compiler Optimization Pass Ordering, 2008.
[3] S. Makula: Compiler Optimization Pass Sequence Exploration using Structured Bayesian Optimization. MSc Dissertation, University of Cambridge, 2017.

[4] Deb, K., Pratap, A., Agarwal, S., & Meyarivan, T. : A Fast Elitist Multi-Objective Genetic Algorithm for Multi-Objective Optimization: NSGA-II. IEEE TEC, 2002.

[5] Van Moffaert, K., & Nowe, A. : Multi-Objective Reinforcement Learning using Sets of Pareto Dominating Policies. JMLR, 2014.

 

10.  Structured Bayesian Optimisation in BoTorch

 

Contact: Eiko Yoneki

 

Optimising system performance is challenging due to the high computational cost of performance evaluation, which is time-consuming and involves navigating a vast search space of variables. Traditional techniques like Bayesian Optimisation (BO) [2] often take a long time to converge and struggle with high-dimensional parameter spaces. However, incorporating structural information into surrogate models can significantly accelerate BO s convergence.

 

This project explores DagBO [1][3], an open-source extension of BO developed by our group, which allows for user-definable surrogate models based on directed acyclic graphs (DAGs). The project will investigate DagBO s performance in tuning Spark benchmarks compared to traditional BO. Potential extensions of the project include distributed computing and the development of sub-models within DagBO.

 

Desirable skills: Computer Systems, Bayesian Optimisation, Spark

 

References:

[1] Ross Tooley: Auto-tuning Spark with Bayesian optimisation.

[2] Jonas B. Mockus. The bayesian approach to global optimization. Freie Univ., Fachbereich

Mathematik, 1984.

[3] https://github.com/Tyv217/dagbo/

 

The project below is a little dated, but if it interests you, you’re very welcome to pick it up and take it further!

11.   Tensor expression superoptimisation via deep reinforcement learning  

Contact: Eiko Yoneki

AlphaDev [1][2] shows RL agent can discover faster sorting algorithm via playing the assembly game. In recent advances of machine learning compiler, EinNet [3] [4] shows how to discover faster tensor programs via rule-based transformation. We are interested in investigating how RL works in discovering faster tensor programs at tensor expression level transformation. You should implement a graph RL agent to play in the tensor game to discover faster tensor programs. Alternatively, we have an internal RL-driven graph transformation system, which you can leverage and apply to the tensor expression level transformation. To get started, build and benchmark EinNet to see how it works with rule-based transformation on a simple DNN (e.g. self-attention). Then, replace their search algorithm with RL algorithm.

References:
[1] Faster sorting algorithms discovered using deep reinforcement learning, Nature, 2023.
[2] https://github.com/google-deepmind/alphadev.
[3] EINNET: Optimizing Tensor Programs with Derivation-Based Transformations. OSDI, 2023.
[4] 
https://github.com/InfiniTensor/InfiniTensor.
[5] 
https://github.com/ucamrl/xrlflow/tree/main.


12.  Better sharding strategy search with deep reinforcement learning

 

Contact: Eiko Yoneki

 

Deep learning recommender model (DLRMs) is one of the most important applications of deep learning. The challenge of DLRMs is to shard the embedding table across multiple devices. This involves column-wise and row-wise sharding. Neuroshard [1] proposes to use a DNN as cost model to guide the search of a sharding strategy, and then it uses a combination of beam search and greedy search to find the sharding strategy. We would like to see whether deep reinforcement learning (RL) can search for a better sharding strategy. Your solution should be compared and benchmarked with [1][2].

 

Desirable skills: Python, deep reinforcement learning

 

References:

[1] Pre-train and Search: Efficient Embedding Table Sharding with Pre-trained Neural Cost Models

[2] AutoShard: Automated Embedding Table Sharding for Recommender Systems

 

13.  Bayesian Optimisation for tensor code generation   

Contact: Eiko Yoneki

Tensor codes are run on massively parallel hardware. When generating tensor code (also called auto-scheduling), TVM Compiler [1] needs to search through many parameters. A State-of-the-art auto-scheduler is Ansor [2], which applies rule-based mutation to generate tensor code templates and then fine-tunes those templates via Evolutionary Search. We think Bayesian Optimisation (BayesOpt) [3] is a better approach to efficiently search the tensor code templates than Evolutionary Search.

At first, TVM is set up for benchmarking with ResNet and Bert using CPU and GPU (possible a few different types of GPU). Next same benchmarking with NVIDIA s Compiler Cutlass should be experimented.

Afterwards, exploring using BayesOpt for high-performance tensor code generation, and benchmark black-box algorithms for tensor code generation. . The main interface for tensor code generation in TVM will be through MetaScheduler [6], which provides a fairly simple Python interface for various search methodologies [7]. We also have a particular interest in tensor code generation for tensor cores, which are equipped by recent generation of GPUs (since the Turing micro-architectures) as a domain-specific architectures to massively accelerate tensor programs.

The project can an advantage of former student s work [9][10], and set the focus on the performance improvement on GPU, Multi-objective BO, and scalability

Desired Skills: Strong interest in tensor program optimisation, Some knowledge/interest in Bayesian optimisation, Python, with some knowledge in C++.

References:
[1] TVM: An Automated End-to-End Optimizing Compiler for Deep Learning https://www.usenix.org/system/files/osdi18-chen.pdf.
[2] Ansor: Generating High-Performance Tensor Programs for Deep Learning https://www.usenix.org/system/files/osdi20-zheng.pdf.
[3] A Tutorial on Bayesian Optimization https://arxiv.org/abs/1807.02811.
[4] HEBO Pushing The Limits of Sample-Efficient Hyperparameter Optimisation https://arxiv.org/abs/2012.03826.
[5] Are Random Decompositions all we need in High Dimensional Bayesian Optimisation? https://arxiv.org/abs/2301.12844.
[6] Tensor Program Optimization with Probabilistic Programs https://arxiv.org/abs/2205.13603.
[7] https://github.com/apache/tvm/blob/4267fbf6a173cd742acb293fab4f77693dc4b887/python/tvm/meta_schedule/search_strategy/search_strategy.py#L238.
[8] NVIDIA Compiler https://github.com/NVIDIA/cutlass.

[9] https://github.com/hgl71964/unity-tvm

[10] Discovering Performant Tensor Programs with Bayesian Optimization

 

14.  Advancing Computer Systems Optimisation through Model-Based Reinforcement Learning 

Contact: Eiko Yoneki


Reinforcement Learning (RL) is gaining interest as a generic optimisation and control method in data management tasks such as resource management, scheduling, database tuning, or stream processing. However, its application is hindered by sample inefficiency and extensive decision evaluation times. Model-Based Reinforcement Learning (MBRL), employing techniques like Probabilistic Ensemble-based models [1], World Models [2] and Dyna-style planning [4] etc. addresses these issues by learning environment models. This project tackles the LLVM compiler optimisation on selection of the pass list. This is critical as pass list selection affects code performance and size, and MBRL s capability to model these interactions is invaluable. Additionally, MBRL challenges on enhancing generalisation ability across different programs, essential due to the expensive training of RL agents. You would use MBRL for intelligent pass list selection in LLVM and work on enhancing generalisation across programs. They will investigate various MBRL techniques and evaluate their effectiveness in modelling the compiler environment and improving transfer learning. You would also evaluate a world-model based approach against a model free approach using software infrastructure. The previous project work [3] can be used as a starting point. Evaluating the project with SPEC 2017 (https://www.spec.org/cpu2017/). benchmarking will be ideal.

Desired Skills: Reinforcement Learning, LLVM Optimisation

References:
[1] Kurtland Chua, Roberto Calandra, Rowan McAllister, and Sergey Levine. 2018. Deep reinforcement learning in a handful of trials using probabilistic dynamics models. NIPS 18. https://dl.acm.org/doi/10.5555/3327345.3327385.
[2] 
World Models. https://arxiv.org/abs/1803.10122.
[3] Yilin Sun: Optimizing LLVM Pass List using Reinforcement Learning.
[4] Sutton, Richard S., et al. Dyna-style planning with linear function approximation and prioritized sweeping. https://arxiv.org/pdf/1206.3285.pdf.

15.   ML-compiler: Joint optimisation of tensor graph and GPU kernels   

Contact: Eiko Yoneki

TASO [2] optimises tensor graph structure, and its backend uses CuDNN to execute tensor Operators [1]. A recent hardware library, Cutlass, allows users to customise their own kernels, as if users can configure their hardware. This brings the opportunities of jointly optimisation of tensor graph and GPU kernels. The project aims to replace TASO s backend from CuDNN [3] to Cutlass [4] . Ideal candidates should have some knowledge writing C++ level codes and preferably know the basics of CUDA programming. The candidates can get started by benchmarking common tensor operators, such as MatMul, Conv2D etc, and compare the performance between different ML compilers. It is a great fit for those who want to know more and deeper about ML compilers.

[1] End-to-end comparison.
[2] Z. Jia, et al.: TASO: Optimizing Deep Learning Computation with Automatic Generation of Graph Substitutions, SOSP, 2019.
[3] CuDNN: NVIDIA cuDNN Installation Guide
[4] Cutlass: CUDA C++ template abstractions

16.   Cost modelling for tensor programs  

Contact: Eiko Yoneki

Cost models provide cheap estimates to get the performance of tensor programs without actual execution, so it is at the heart of accelerating tenor programs generation. For example, TVM [1] uses XGboost [2] to estimate tensor program runtime. More recently, advanced DNN-based cost models are proposed to perform more accurate estimation such as GNN [3] and transformers [4]. Cost modelling can be even cross-hardware [5]. On the other hand, distributed DNN training also needs delicate cost modelling to provide cheap estimation of the throughput of parallelisation strategies, such as [6] and [7]. However, distributed cost modelling is mostly mathematical estimation. This project aims to investigate and benchmark the state-of-the-art cost modelling. In particular, we want to understand how good those cost models are, and if possible, we want to replace the distributed cost modelling with a DNN-based cost model.

[1] TVM: An Automated End-to-End Optimizing Compiler for Deep Learning.
[2] AutoTVM: Learning to Optimize Tensor Programs.
[3] A Graph Neural Network-Based Performance Model for Deep Learning Applications.

Contact Email

Please email to eiko.yoneki@cl.cam.ac.uk for any question.