AI Briefing
KO

Hugging Face Announces Open-Source Dataset Achievements

·2024.06.20 09:00

Key point

Hugging Face unveiled the dataset-building achievements and future plans of the collaborative project 'Data Is Better Together'.

Details

Hugging Face announced key achievements and a future roadmap for the Data Is Better Together (DIBT) initiative, a collaboration between Hugging Face and Argilla.

Key Achievements

  • Prompt Ranking Project: Released the DIBT/10k_prompts_ranked dataset, containing quality rankings for 10,000 prompts. This dataset is being used to build new models such as SPIN.
  • Multilingual Prompt Evaluation Project (MPEP): Developing a multilingual prompt evaluation benchmark to address English-centric data bias. Translation work is currently underway for multiple languages including Dutch, Russian, and Spanish.

Future Plans (Cookbook efforts) To enable the community to build valuable datasets on their own, the following guides and tools will be supported:

  • Domain-Specific Datasets: Connecting engineers with domain experts to support the building of datasets for specific fields.
  • DPO/ORPO and KTO Datasets: Providing tools to help build DPO-style and KTO datasets tailored to various languages and tasks.

Hugging Face aims to expand its dataset builder community through Discord and resolve data imbalance in the open-source ecosystem.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.