Skip to main content
Too Many; Didn't Read? Classifying Large Datasets with LLMs
TalksFree

Too Many; Didn't Read? Classifying Large Datasets with LLMs

Similar events

View all →

About This Event

The event will focus on using large language models (LLMs) to classify datasets and extract information at scale, while also addressing the reliability of the results. It takes place on 26 October 2026 at Cambridge Digital Humanities.

Using large language models (LLMs) to classify datasets and extract information at scale, and on knowing when the results can be trusted

Convenor

Raphael Hernandes (https://www.cdh.cam.ac.uk/about/people/raphael-hernandes/?category=1946)

Raphael is a PhD candidate at Cambridge Digital Humanities, researching how AI impacts journalism, information environments, and mediation. His academic work comes after more than a decade as a journalist at places like Folha de S.Paulo and the Guardian.

Description

This workshop covers how to use large language models (LLMs) to classify datasets and extract information at scale, and on knowing when the results can be trusted. Researchers often face corpora too large to read and classify manually. Data from social media and other online platforms pose a greater challenge due to chaotic communication, slang, and shifting meanings across communities. Fixed dictionaries, keyword counts, and older machine learning models can be tricky to implement and often fail at complex labelling in this kind of data. LLMs and multimodal models can code (in the social sciences sense) and extract information at scale. Outputs, however, are probabilistic by design; without testing, there is no way to know whether a classification is reliable. Thus, validation is an essential step. The workshop moves in four parts. First, context: where LLMs are appropriate against manual coding or supervised machine learning. Second, the hands-on core: designing a codebook, writing structured prompts, classifying a real dataset. Third, validation: building gold-standard subsets, measuring agreement, and analysing errors. Fourth, extensions: cross-checking across models and extracting information from images. No programming experience is necessary for this workshop, though Python notebooks will be presented as an option for those who prefer them.

Is any equipment or Software required?

Participants need a laptop, a web browser, and a Google Account. The workshop uses Google's Gemini API, which has a free tier that requires no credit card or payment details. Participants create their own API key through Google AI Studio using a personal Google account. Setup instructions will be circulated in advance. Both tracks run in the browser, with nothing to install: - No-code: a Google Sheets template I provide, which calls the model from a spreadsheet formula. - Code: a Google Colab notebook I provide. Exercises are sized to stay within free-tier rate limits. I will also supply pre-computed model outputs so that the validation exercises — the core of the session — can be completed with no API access at all, should keys or connectivity fail. The methods are not tied to any one provider. The same workflow applies to other models, and the materials note alternatives, including locally-run ones.

Target Audience

Our CDH Methods workshops have limited places and are prioritised for students and staff at the University of Cambridge. However, if space is available, we welcome all participants who want to learn and apply digital methods and use digital tools in their research.

This session may be of particular interest to:

PhD students in the Arts, Humanities and Social Sciences

Early Career Researchers in the Arts, Humanities and Social Sciences

Contact CDH

If you have specific accessibility needs for this event, please get in touch. We will do our best to accommodate any requests, however please note that the building is grade 2 listed and has 4 steps into the building so unfortunately wheelchair access is not available.

This workshop is part of our Methods Fellowship programme, which develops and delivers innovative teaching in digital methods. You can read more about the programme here (https://www.cdh.cam.ac.uk/methods/methods-fellowships-2026-27/)and view the complete series of workshops here (https://www.cdh.cam.ac.uk/events/?type=1813&range=upcoming).

Description supplied by the event organiser.

Staying overnight?

Hotels within walking distance of Cambridge Digital Humanities, 2 Trumpington Street, Cambridge.

Searching near 2 Trumpington Street, Cambridge, CB2 1QA
Check-in: Monday, 26 October 2026
Find hotels near this venuecompare prices →

Ad · Stay22 affiliate link

Similar Events in Cambridge

View all →

More at Cambridge Digital Humanities, 2 Trumpington Street, Cambridge

View all →

Frequently asked questions

Too Many; Didn't Read? Classifying Large Datasets with LLMs takes place on Monday, 26 October 2026 at 13:00.

Too Many; Didn't Read? Classifying Large Datasets with LLMs is held at Cambridge Digital Humanities, 2 Trumpington Street, Cambridge in Cambridge (2 Trumpington Street, Cambridge, CB2 1QA).

Entry to Too Many; Didn't Read? Classifying Large Datasets with LLMs is free.

Yes — hotels within walking distance of Cambridge Digital Humanities, 2 Trumpington Street, Cambridge can be compared using the accommodation search further down this page.

Tickets for Too Many; Didn't Read? Classifying Large Datasets with LLMs can be booked via Eventbrite using the "Check Tickets & Live Prices" button on this page, which opens the official booking site in a new tab.

Price

Free Entry

Get Tickets