Introduction to Computational Semantics and Pragmatics
Principal lecturer: Prof Simone Teufel
Taken by: MPhil ACS, Part III
Code: L99
Term: Michaelmas
Hours: 16
Format: In-person (lectures, pre-lecture reading, annotation of texts and discussion)
Class limit: max. 16 students
timetable
Aims
This course is an overview course of theoretical semantics, computational semantics, inference, and pragmatics. It is a theoretical course with some practical exercises. It covers truth-conditional logic-based formalisms and other meaning representations, including the principle of semantic compositionality and scope. Based on these representations for the meaning of individual sentences, the course also explores a range of semantic phenomena that extend beyond the sentence level. These include negation, some aspects of pragmatics, and discourse structure. This course is relevant in the light of interpretation of LLM output. LLMs are clearly able to produce grammatically fluent and thematically relevant output, but it is far more difficult to assess whether this output is also semantically appropriate. One of the tools that enable students to perform such objective assessment is a close study of the semantic properties of human language. Amongst other skills, students will learn to distinguish relevant semantic and pragmatic phenomena, find ambiguity in text and identify cases of distorted meaning. They will also learn how meaning emerges from phenomena above the sentence level. Knowledge of these aspects of semantics of language will provide a solid foundation for meaningful evaluation of modern LLM and other automatically produced text.
Format
The course is lecture-based with weekly reading, in 8 sessions. Discussion sessions follow each homework (3x).
Practical advantages of this course for NLP students
Knowledge of underlying semantic effects helps improve NLP evaluation, for instance by providing more meaningful error analysis. You will be able to link particular errors to design decisions inside your system. You will learn methods for better benchmarking of your system, whatever the task may be. Supervised ML systems (in particular black-box systems such as Deep Learning) are only as clever as the datasets they are based on. In this course, you will learn to critique existing datasets, and to design datasets so that they are harder to trick without real understanding. You will be able to design tests for ML systems that better pinpoint which aspects of language an end-to-end system has “understood”. You will also learn to detect ambiguity and ill-formed semantics in human-human communication. This can serve to improve your scientific writing, by planning text more clearly and logically.
Topics
Events and semantic role labelling
Referentiality and coreference resolutions
Truth-conditional semantics
Vector space models
Compositional semantics
Negation and scope
Information Status in Discourse
Entailment
Inference
Presupposition and Gricean pragmatics
Assessment
20% Coursework, to be delivered in 1 unassessed and 2 assessed homeworks (week 1, 3 and 6):
- Homework 1:feedback will be provided but not assessed
- Homework 2: 20%
- Homework 3: 20%
60% In-class test (last session)
The test will check basic knowledge of concepts taught such as terminology. Questions may be presented in multiple-choice format or require longer written responses. Questions may also be research-related, relating to readings that have been assigned during class or specifically as preparation for the test. Students may be asked to provide their own examples of linguistic phenomena, interpret given examples, and carry out mathematical derivations. To help students prepare for the exam and to make clear which material is examinable, each lecture will highlight possible exam questions in the form of core 'take-home' knowledge and thinking-further questions which students can solve in their own time after the lecture.
Reading
Session 1: He, Lewis, Zettlemoyer (2015): Question-Answer Driven Semantic Role Labeling: Using Natural Language to Annotate Natural Language. EMNLP 2015. pdf here
Session 2: Pradhan, Moschitti, Xue, Uryupina, Zhang (2012): CoNLL-2012 Shared Task: Modeling Multilingual Unrestricted Coreference in OntoNotes. Joint Conference on EMNLP and CoNLL. pdf here (Read up to 4.4.2; throughout ignore Chinese and Arabic)
Additional reading to be added 1 week before each session