AI Briefing
KO

Compressing All of Humanity's Recipes into 2MB

·2026.05.27 23:14

Key point

The paper proposes 'Epicure,' an embedding model trained on 4.14 million multilingual recipes that reflects the chemical properties of ingredients and their culinary context.

Details

Epicure is a family of ingredient embedding models trained on 4.14 million recipe data points spanning 11 languages.

Through an LLM-based pipeline, the vast ingredient data from recipes was normalized into 1,790 standard items, and the model leverages a graph structure that combines ingredient co-occurrence relationships (NPMI graph) with chemical compound information (FlavorDB).

The researchers analyze ingredient relationships from multiple angles through three variants of Metapath2Vec:

  • Cooc: Learns a graph of simple co-occurrence relationships between ingredients
  • Chem: Learns chemical compound metapaths
  • Core: Combines co-occurrence relationships and chemical properties to balance between the two domains

Through this, the model can finely embed the relationship between the chemical properties of ingredients and their actual culinary context.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.