• English
  • Deutsch
  • Log In
    Password Login
    Research Outputs
    Fundings & Projects
    Researchers
    Institutes
    Statistics
Repository logo
Fraunhofer-Gesellschaft
  1. Home
  2. Fraunhofer-Gesellschaft
  3. Anderes
  4. Data Processing for the OpenGPT-X Model Family
 
  • Details
  • Full
Options
October 11, 2024
Paper (Preprint, Research Paper, Review Paper, White Paper, etc.)
Title

Data Processing for the OpenGPT-X Model Family

Title Supplement
Published on arXiv
Abstract
This paper presents a comprehensive overview of the data preparation pipeline developed for the OpenGPT-X project, a large-scale initiative aimed at creating open and high-performance multilingual large language models (LLMs). The project goal is to deliver models that cover all major European languages, with a particular focus on real-world applications within the European Union. We explain all data processing steps, starting with the data selection and requirement definition to the preparation of the final datasets for model training. We distinguish between curated data and web data, as each of these categories is handled by distinct pipelines, with curated data undergoing minimal filtering and web data requiring extensive filtering and deduplication. This distinction guided the development of specialized algorithmic solutions for both pipelines. In addition to describing the processing methodologies, we provide an in-depth analysis of the datasets, increasing transparency and alignment with European data regulations. Finally, we share key insights and challenges faced during the project, offering recommendations for future endeavors in large-scale multilingual data preparation for LLMs.
Author(s)
Brandizzi, Nicolo  
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Abdelwahab, Hammam
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Bhowmick, Anirban
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Stein, Benny Jörg  orcid-logo
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Denisov, Pavel
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Helmer, Lennard  orcid-logo
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Fromm, Michael  
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Saleem, Qasid
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Ali, Mehdi  
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Rutmann, Richard
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Naderi, Farzad
Fraunhofer-Institut für Integrierte Schaltungen IIS  
Agy, Mohamad Saif
Fraunhofer-Institut für Integrierte Schaltungen IIS  
Schwirjow, Alexander
Fraunhofer-Institut für Integrierte Schaltungen IIS  
Küch, Fabian  
Fraunhofer-Institut für Integrierte Schaltungen IIS  
Hahn, Luzian
Fraunhofer-Institut für Integrierte Schaltungen IIS  
Ostendorff, Malte  
DFKI  
Ortiz Suarez, Pedro
DFKI  
Rehm, Georg  
DFKI  
Wegener, Dennis  orcid-logo
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Flores-Herr, Nicolas  
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Köhler, Joachim  
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Leveling, Johannes
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Project(s)
Aufbau eines Gaia-X Knotens für große KI-Sprachmodelle und innovative Sprachapplikations-Services  
Funder
Bundesministerium für Wirtschaft und Klimaschutz  
DOI
10.48550/arXiv.2410.08800
Language
English
Fraunhofer-Institut für Intelligente Analyse- und Informationssysteme IAIS  
Fraunhofer-Institut für Integrierte Schaltungen IIS  
Keyword(s)
  • NLP

  • Data

  • OpenGPT-X

  • LLM

  • Cookie settings
  • Imprint
  • Privacy policy
  • Api
  • Contact
© 2024