The course introduces natural language processing (NLP) and large language models (LLMs) to students from humanities, social sciences, and interdisciplinary programmes. It explains the foundations of traditional, statistical and neural approaches and reinforces them through hands-on use of open-source systems, primarily in Python. A substantial part of the course is devoted to working with generative LLMs (e.g. ChatGPT, Claude, open models such as LLaMA, Mistral, Gemma) and to prompt engineering techniques. During tutorials students solve realistic problems that involve processing large volumes of textual data, including resources in Slovene.
Topics of lectures and tutorials:
- Introduction to NLP and LLMs: what natural language is from a computer's perspective, ambiguity and linguistic complexity, a brief history of NLP, from traditional to neural approaches, where LLMs appear in everyday use.
- Foundations of working with text: text preparation and cleaning, character encodings, optical character recognition (OCR) and transcription, language normalisation, lemmatisation and morphological tagging, regular expressions, tokenisation (including modern approaches – e.g., BPE and WordPiece).
- Similarity measures and basic text representations: bag-of-words, TF-IDF, cosine similarity, n-grams, text clustering, simple sentiment classification.
- Language resources: corpora (Gigafida, ccGigafida, slWaC, Common Crawl, Wikipedia), annotated resources (SUK, benchmarking corpora), repositories (CLARIN.SI, Hugging Face) and European Language Data Space, licensing and legal aspects, FAIR principles, ethics of data collection.
- Deep neural networks and word embeddings: an intuitive explanation of neural networks, contextual and non-contextual embeddings (word2vec, fastText), multilingual embeddings, sentence and document similarity.
- Large language models: the transformer architecture (at an intuitive level), encoder and/or decoder families (BERT, GPT, T5), how LLMs are trained, reinforcement learning from human feedback (RLHF), open vs. closed models, running models locally (Ollama), cost, speed and environmental aspects, LLMs as a service.
- Prompt engineering: fundamentals of writing instructions, techniques (zero-shot, few-shot, chain-of-thought), structured outputs (JSON), system prompts and roles, common pitfalls and hallucinations, recipes for practical tasks (classification, summarisation, extraction, translation, text editing).
- Retrieval-augmented generation (RAG): document preparation, vector databases (FAISS, Chroma), search engines and text retrieval, building a RAG over your own documents.
- Semantic and lexical technologies: WordNet and sloWNet, FrameNet, BabelNet, ontologies and knowledge graphs (WikiData), Wikidata, entity linking, terminological resources; how to complement LLMs with structured knowledge.
- Linguistic tasks and multilinguality: named entity recognition, translation, word sense disambiguation, sentiment analysis, coreference resolution; specifics of Slovene and other morphologically rich languages.
- Evaluation and reliability of LLMs: automatic metrics (BLEU, ROUGE, BERTScore), human evaluation, LLM-as-a-judge, evaluation frameworks (SloBENCH), bias, safety and ethical considerations.
- Agentic technologies and tools: function calling, integration with external resources, the MCP paradigm, agents that assist with research, writing and data analysis.
- Language technologies in practice: case studies from humanistics and social sciences - translation studies, linguistic research, social sciences, education, librarianship and history; designing the student's own mini-project.
- Multimodal and speech models (brief overview): automatic speech recognition (e.g. Whisper), text-to-speech, text-to-image models, intuitive overview for tasks that go beyond text only.
Project work: in small groups or individually, students tackle a chosen problem on a larger volume of textual data (e.g. analysis of archival material/books, building a chatbot over a corpus, comparative analysis of models, processing social-media language data).