
Too Many; Didn't Read? Classifying Large Datasets with LLMs
Similar events
View all →About This Event
The event will focus on using large language models (LLMs) to classify datasets and extract information at scale, while also addressing the reliability of the results. It takes place on 26 October 2026 at Cambridge Digital Humanities.
Using large language models (LLMs) to classify datasets and extract information at scale, and on knowing when the results can be trusted
Convenor
Raphael Hernandes (https://www.cdh.cam.ac.uk/about/people/raphael-hernandes/?category=1946)
Raphael is a PhD candidate at Cambridge Digital Humanities, researching how AI impacts journalism, information environments, and mediation. His academic work comes after more than a decade as a journalist at places like Folha de S.Paulo and the Guardian.
Description
This workshop covers how to use large language models (LLMs) to classify datasets and extract information at scale, and on knowing when the results can be trusted. Researchers often face corpora too large to read and classify manually. Data from social media and other online platforms pose a greater challenge due to chaotic communication, slang, and shifting meanings across communities. Fixed dictionaries, keyword counts, and older machine learning models can be tricky to implement and often fail at complex labelling in this kind of data. LLMs and multimodal models can code (in the social sciences sense) and extract information at scale. Outputs, however, are probabilistic by design; without testing, there is no way to know whether a classification is reliable. Thus, validation is an essential step. The workshop moves in four parts. First, context: where LLMs are appropriate against manual coding or supervised machine learning. Second, the hands-on core: designing a codebook, writing structured prompts, classifying a real dataset. Third, validation: building gold-standard subsets, measuring agreement, and analysing errors. Fourth, extensions: cross-checking across models and extracting information from images. No programming experience is necessary for this workshop, though Python notebooks will be presented as an option for those who prefer them.
Is any equipment or Software required?
Participants need a laptop, a web browser, and a Google Account. The workshop uses Google's Gemini API, which has a free tier that requires no credit card or payment details. Participants create their own API key through Google AI Studio using a personal Google account. Setup instructions will be circulated in advance. Both tracks run in the browser, with nothing to install: - No-code: a Google Sheets template I provide, which calls the model from a spreadsheet formula. - Code: a Google Colab notebook I provide. Exercises are sized to stay within free-tier rate limits. I will also supply pre-computed model outputs so that the validation exercises — the core of the session — can be completed with no API access at all, should keys or connectivity fail. The methods are not tied to any one provider. The same workflow applies to other models, and the materials note alternatives, including locally-run ones.
Target Audience
Our CDH Methods workshops have limited places and are prioritised for students and staff at the University of Cambridge. However, if space is available, we welcome all participants who want to learn and apply digital methods and use digital tools in their research.
This session may be of particular interest to:
PhD students in the Arts, Humanities and Social Sciences
Early Career Researchers in the Arts, Humanities and Social Sciences
Contact CDH
If you have specific accessibility needs for this event, please get in touch. We will do our best to accommodate any requests, however please note that the building is grade 2 listed and has 4 steps into the building so unfortunately wheelchair access is not available.
This workshop is part of our Methods Fellowship programme, which develops and delivers innovative teaching in digital methods. You can read more about the programme here (https://www.cdh.cam.ac.uk/methods/methods-fellowships-2026-27/)and view the complete series of workshops here (https://www.cdh.cam.ac.uk/events/?type=1813&range=upcoming).
Description supplied by the event organiser.
Staying overnight?
Hotels within walking distance of Cambridge Digital Humanities, 2 Trumpington Street, Cambridge.
Ad · Stay22 affiliate link
Similar Events in Cambridge
View all →
From £5An Evening with Rachel Harrison in Cambridge
Fri, 2 Oct 2026
Waterstones
£0.00Removing gas on a budget, Mortlock Avenue
Sat, 3 Oct 2026
Mortlock Avenue
£0.00Grant funded full retrofit, Fernlea Close, Cherry Hinton
Sat, 3 Oct 2026
Fernlea Close
Free EntryBhagavad Gita for Hard Times
Sat, 3 Oct 2026
Arbury Community Centre
Free EntryWho Presses the Button? Human Oversight of AI Between Claim and Practice
Mon, 5 Oct 2026
Nokia Bell Labs Cambridge
£16.96Present With Confidence
Mon, 5 Oct 2026
St Ives Corn Exchange
More at Cambridge Digital Humanities, 2 Trumpington Street, Cambridge
View all →Frequently asked questions
Too Many; Didn't Read? Classifying Large Datasets with LLMs takes place on Monday, 26 October 2026 at 13:00.
Too Many; Didn't Read? Classifying Large Datasets with LLMs is held at Cambridge Digital Humanities, 2 Trumpington Street, Cambridge in Cambridge (2 Trumpington Street, Cambridge, CB2 1QA).
Entry to Too Many; Didn't Read? Classifying Large Datasets with LLMs is free.
Yes — hotels within walking distance of Cambridge Digital Humanities, 2 Trumpington Street, Cambridge can be compared using the accommodation search further down this page.
Tickets for Too Many; Didn't Read? Classifying Large Datasets with LLMs can be booked via Eventbrite using the "Check Tickets & Live Prices" button on this page, which opens the official booking site in a new tab.

Too Many; Didn't Read? Classifying Large Datasets with LLMs
Editorial Summary
The event will focus on using large language models (LLMs) to classify datasets and extract information at scale, while also addressing the reliability of the results. It takes place on 26 October 2026 at Cambridge Digital Humanities.
This summary was automatically generated from event data sourced from official ticketing providers.
Organiser’s full description
Using large language models (LLMs) to classify datasets and extract information at scale, and on knowing when the results can be trusted
Convenor
Raphael Hernandes (https://www.cdh.cam.ac.uk/about/people/raphael-hernandes/?category=1946)
Raphael is a PhD candidate at Cambridge Digital Humanities, researching how AI impacts journalism, information environments, and mediation. His academic work comes after more than a decade as a journalist at places like Folha de S.Paulo and the Guardian.
Description
This workshop covers how to use large language models (LLMs) to classify datasets and extract information at scale, and on knowing when the results can be trusted. Researchers often face corpora too large to read and classify manually. Data from social media and other online platforms pose a greater challenge due to chaotic communication, slang, and shifting meanings across communities. Fixed dictionaries, keyword counts, and older machine learning models can be tricky to implement and often fail at complex labelling in this kind of data. LLMs and multimodal models can code (in the social sciences sense) and extract information at scale. Outputs, however, are probabilistic by design; without testing, there is no way to know whether a classification is reliable. Thus, validation is an essential step. The workshop moves in four parts. First, context: where LLMs are appropriate against manual coding or supervised machine learning. Second, the hands-on core: designing a codebook, writing structured prompts, classifying a real dataset. Third, validation: building gold-standard subsets, measuring agreement, and analysing errors. Fourth, extensions: cross-checking across models and extracting information from images. No programming experience is necessary for this workshop, though Python notebooks will be presented as an option for those who prefer them.
Is any equipment or Software required?
Participants need a laptop, a web browser, and a Google Account. The workshop uses Google's Gemini API, which has a free tier that requires no credit card or payment details. Participants create their own API key through Google AI Studio using a personal Google account. Setup instructions will be circulated in advance. Both tracks run in the browser, with nothing to install: - No-code: a Google Sheets template I provide, which calls the model from a spreadsheet formula. - Code: a Google Colab notebook I provide. Exercises are sized to stay within free-tier rate limits. I will also supply pre-computed model outputs so that the validation exercises — the core of the session — can be completed with no API access at all, should keys or connectivity fail. The methods are not tied to any one provider. The same workflow applies to other models, and the materials note alternatives, including locally-run ones.
Target Audience
Our CDH Methods workshops have limited places and are prioritised for students and staff at the University of Cambridge. However, if space is available, we welcome all participants who want to learn and apply digital methods and use digital tools in their research.
This session may be of particular interest to:
PhD students in the Arts, Humanities and Social Sciences
Early Career Researchers in the Arts, Humanities and Social Sciences
Contact CDH
If you have specific accessibility needs for this event, please get in touch. We will do our best to accommodate any requests, however please note that the building is grade 2 listed and has 4 steps into the building so unfortunately wheelchair access is not available.
This workshop is part of our Methods Fellowship programme, which develops and delivers innovative teaching in digital methods. You can read more about the programme here (https://www.cdh.cam.ac.uk/methods/methods-fellowships-2026-27/)and view the complete series of workshops here (https://www.cdh.cam.ac.uk/events/?type=1813&range=upcoming).
Description supplied by the event organiser.
📍 Venue
Cambridge Digital Humanities, 2 Trumpington Street, Cambridge2 Trumpington Street, Cambridge, CB2 1QA
Market, Cambridge
Get Directions →❓ Frequently asked questions
When is Too Many; Didn't Read? Classifying Large Datasets with LLMs?
Too Many; Didn't Read? Classifying Large Datasets with LLMs takes place on Monday, 26 October 2026 at 13:00.
Where is Too Many; Didn't Read? Classifying Large Datasets with LLMs taking place?
Too Many; Didn't Read? Classifying Large Datasets with LLMs is held at Cambridge Digital Humanities, 2 Trumpington Street, Cambridge in Cambridge (2 Trumpington Street, Cambridge, CB2 1QA).
How much are tickets for Too Many; Didn't Read? Classifying Large Datasets with LLMs?
Entry to Too Many; Didn't Read? Classifying Large Datasets with LLMs is free.
Is there anywhere to stay near Too Many; Didn't Read? Classifying Large Datasets with LLMs?
Yes — hotels within walking distance of Cambridge Digital Humanities, 2 Trumpington Street, Cambridge can be compared using the accommodation search further down this page.
How do I get tickets for Too Many; Didn't Read? Classifying Large Datasets with LLMs?
Tickets for Too Many; Didn't Read? Classifying Large Datasets with LLMs can be booked via Eventbrite using the "Check Tickets & Live Prices" button on this page, which opens the official booking site in a new tab.
🏨 Staying overnight?
Hotels within walking distance of Cambridge Digital Humanities, 2 Trumpington Street, Cambridge
Ad · Stay22 affiliate link
Price
Free Entry
Date
Until Monday, 26 October 2026
Time
Venue
2 Trumpington Street, Cambridge, CB2 1QA
More in Cambridge
EventsList curates 183,000+ events across the UK. Tickets purchased via the venue or authorised sellers only.
More Events at Cambridge Digital Humanities, 2 Trumpington Street, Cambridge
View all →
FREEHow Does GenAI Categorise You?
📍 Cambridge Digital Humanities, 2 Trumpington Street, Cambridge
FREEMapping Historical Networks
📍 Cambridge Digital Humanities, 2 Trumpington Street, Cambridge
FREEWhen More Talking isn't More Learning
📍 Cambridge Digital Humanities, 2 Trumpington Street, Cambridge
FREEBefore the Answer: Visual Methods and the Preservation of Judgement
📍 Cambridge Digital Humanities, 2 Trumpington Street, Cambridge
Similar Events in Cambridge
View all →
An Evening with Rachel Harrison in Cambridge
📍 Waterstones
£5.00 - £13.00

Removing gas on a budget, Mortlock Avenue
📍 Mortlock Avenue
£0.00

Grant funded full retrofit, Fernlea Close, Cherry Hinton
📍 Fernlea Close
£0.00
FREEBhagavad Gita for Hard Times
📍 Arbury Community Centre
FREEWho Presses the Button? Human Oversight of AI Between Claim and Practice
📍 Nokia Bell Labs Cambridge
SOLD OUTPresent With Confidence
📍 St Ives Corn Exchange
Sold out at source
More things to do in Cambridge
- Key ingredients for nature recovery· Friday, 13 November 2026
- Laura Smyth: Born Aggy· Friday, 9 October 2026
- 2026 - Foro Cervantes at the University of Cambridge with David Uclés· Thursday, 15 October 2026
- Sylvia Plath, Witchcraft & Radical Poetry· Monday, 12 October 2026
- Voices That Shape Research· Friday, 16 October 2026
- 'We Flow On': Robert Macfarlane and Mark Wormald on Living Water· Tuesday, 3 November 2026
- The 2026 Marshall Professor Inaugural Lecture: Pooja Agrawal· Friday, 16 October 2026
- Tea & Talk with Naja Hendriksen: Meeting Contemporary Greenland· Friday, 16 October 2026
- The Organization of Goal-Directed Behaviour in the Brain· Monday, 26 October 2026
- CAHS MT26 - Matthew Walker, Assistant Professor of British Architecture· Monday, 26 October 2026
- Climate shocks, conflict and primary education in Ethiopia· Tuesday, 27 October 2026
- Newmarket Brook: Past Present and Future· Wednesday, 28 October 2026
- Pathways: A Practice Sharing Event with Jane Pryor and Polly Griffiths· Saturday, 24 October 2026
- The Causes of Great-Power War· Friday, 23 October 2026

