Postdoctoral Researcher in Artificial Intelligence and Natural Language Processing (SCAI/BnF research program)
Who we are
Sorbonne University is a multidisciplinary research university established on January 1, 2018, through the merger of Paris-Sorbonne and Pierre and Marie Curie universities. Offering educational programs to 54,000 students, including 4,700 doctoral candidates and 10,200 international students, it employs 6,300 professors, teacher-researchers, and researchers, alongside 4,900 library, administrative, technical, social, and healthcare staff. It has an annual budget of €670 million. Sorbonne University possesses world-class potential, primarily located in the heart of Paris, with a presence spanning over twenty sites in the Île-de-France region and other territories. Sorbonne University features a unique organizational structure across three Faculties—Humanities, Science & Engineering, and Medicine—which enjoy significant autonomy in implementing the university’s strategy within their respective scopes, governed by performance and resources agreements. The university leadership prioritizes promoting university strategy, governance, partnership development, and revenue diversification.
About the Department / Structure
In a national and international landscape shaped by intense competition in artificial intelligence, Sorbonne University established the "Sorbonne Center for Artificial Intelligence" (SCAI). Situated in a single, central location in the Latin Quarter, SCAI brings together a strategic range of modern artificial intelligence disciplines. SCAI’s ambition is to make a significant contribution to excellence in interdisciplinary AI research by fostering collaboration between teacher-researchers, researchers, teachers, students, and industry partners. In 2024, the Sorbonne University alliance was recognized as one of France's leading national centers of excellence in Artificial Intelligence and awarded €35 million in competitive funding to develop the "PostGenerativeAI" project, with SCAI serving as its driving force. Among the Sorbonne University laboratories involved in SCAI, ISIR (Institute of Intelligent Systems and Robotics) brings together research teams active in robotics, cognitive sciences, human-computer interaction, and machine learning.
The research project described below is part of a strategic partnership between Sorbonne University and the National Library of France (BnF). This partnership specifically combines the expertise of the Machine Learning – Deep Learning and Information Access (MLIA) team at the Institute of Intelligent Systems and Robotics (ISIR) with that of the BnF to conduct joint research on large language models (LLMs).
The Bibliothèque nationale de France (BnF) is one of the largest heritage libraries in the world. Its mission is to collect, catalog, preserve, enrich, and share the national documentary heritage. Having been committed for many years to ambitious digitization programs for its collections—now further expanded by the massive intake of born-digital collections—the BnF continuously enriches its digital heritage. The sheer volume, diversity, and rapid growth of this material demand new processing and consulting tools. To enable as many people as possible to discover and engage with this heritage, the BnF has been actively involved in artificial intelligence (AI) technologies for several years.
Project Description
DATA DESCRIPTION
The BnF web archives, preserved as part of its legal deposit mission, offer a unique source for studying the present era. Although rich in content (spanning 30 years, 2.5 PB of data, and 59 billion URLs) and multi-semiotic (text, image, audio, video), they remain under-exploited due to their sheer volume, heterogeneity, and specific formats. Fully indexed by URL and date, but only partially indexed in full-text via Apache Solr, these collections could be better leveraged using deep learning tools, particularly to facilitate the creation and exploration of sub-corpora. The objective of this project is to develop tools and representations (typologies, relational graphs, semantic maps) to help researchers formulate research questions and analyze these massive datasets.
Two main areas of focus have emerged:
-
Analysis of a complete year of the French web: This will involve mapping the entirety of the collection gathered during the year 2004. Beyond the practical challenges of preparing a text corpus from the BnF archives, the focus will be on mapping the archive from the perspective of its linguistic and thematic content, cataloging idioms and their variations, registers, and genres, as well as their associated themes.
-
Electoral web archive: This will involve studying a sub-collection corresponding to archives collected during national elections in all their diversity (press, blogs, institutional communications from parties and unions, social networks) in order to document, for example, the evolution of political communication strategies, the reconfiguration of debates around certain controversial topics, or to build empirical indicators of the diversity of viewpoints.
Key Missions
This project consists of working on the exploitation of web archives using advanced artificial intelligence and text mining methods. Particular attention will be paid to the algorithmic efficiency of the designed and deployed processing pipelines in relation to the volume of data to be processed.
The successful candidate will be responsible for:
-
Implementing NLP (Natural Language Processing) modules for large text corpora derived from BnF archives.
-
Developing algorithms based on machine learning methodologies (classification, representation learning, fine-tuning of large language models).
-
Utilizing, where appropriate, large language models (LLMs) to develop new processing methods, perform semi-automatic annotations, or evaluate existing processing chains.
-
Reporting and presenting development work in a clear and efficient manner, both for discussions with BnF experts and for writing scientific publications and presenting them at specialized conferences or workshops.
The developed methods and tools are intended to enrich the service offerings provided by the BnF.
Qualifications / Education
A PhD in computer science or equivalent is required, along with a strong scientific track record, particularly in Natural Language Processing (NLP) and/or Text Mining and/or Information Retrieval. Experience with international research projects and applications in the Humanities and Social Sciences would be an asset.
General Information
-
Locations: Pierre and Marie Curie Campus of Sorbonne University and the François-Mitterrand site of the BnF.
-
Contract: Fixed-term contract (CDD) of 12 months with the possibility of extension.
-
Expected start date: As soon as possible.
-
Working hours: Full-time.
-
Desired experience: 1 to 3 years.
-
Salary: Commensurate with experience.
Key Contacts
-
François Yvon, CNRS Research Director, Member of the MLIA team at ISIR (Sorbonne University)
-
Dorothée Benhamou-Suesser, Head of Access to Internet Archives (BnF, Digital Legal Deposit Department)
-
Jean-Philippe Moreux, Head of Artificial Intelligence Mission
-
Supervisory / Management role: NO
-
Project management: YES
A solid background in natural language processing or text analysis is essential, and very strong programming skills (Python, PyTorch) are required. An understanding of the ethical issues surrounding such systems is also expected.
-
Language: Knowledge of French is not mandatory but highly desirable.
How to Apply
Applications (CV, letter of motivation, references if available..), should be sent to