Apple Research: Tree-Structured Expression Serialization Limits LLM Communication
Key point
Fine-tuning with ~3,600 examples enables open-weight models to surpass untrained Gemini-3.1-Pro in expression recovery.
Details
Language models lose significant information when serializing tree-structured compositional content into natural language, a process common in chain-of-thought reasoning. Apple researchers propose a round-trip protocol to measure this loss, using a generator to convert arithmetic expressions into word problems and a separate extractor to recover them, with symbolic equivalence serving as the exact oracle.
Asymmetric Communication Loss
The study evaluated all pairwise combinations of 16 models, revealing that the communication channel is both lossy and asymmetric. Swapping which model generates the text and which extracts the expression shifts accuracy by up to 60.4 points. The best performance, reaching 92.9%, occurred when combining different models for generation and extraction rather than using the same model for both tasks.
Structural Drivers of Failure
At least 73.6% of round-trip failures originate during the generation phase. The difficulty is driven by the tree structure of the expression—specifically operator count, depth, and right-branching—rather than the model family. This suggests that the complexity of the hierarchical data itself is the primary bottleneck.
Trainability of the Channel
The serialization bottleneck is trainable. Using approximately 3,600 fine-tuning examples that share the evaluation's operators and tree shapes, every open-weight model tested lifted above the performance of untrained Gemini-3.1-Pro. This result holds even in a disjoint-domain regime with new operators and vocabulary, confirming the gains are not artifacts of matched semantics, though a gap to frontier models remains.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.