Knowledge Distillation of Black-Box Large Language Models
Key point
The paper proposes Proxy-KD, a methodology that leverages a proxy model to efficiently transfer knowledge from black-box LLMs such as GPT-4.
Details
Proprietary LLMs like GPT-4 possess excellent performance, but their Black-box nature—lacking access to internal states—imposes limitations on information transfer during Knowledge Distillation (KD).
To overcome this limitation, this paper introduces a new methodology called Proxy-KD. This approach uses a Proxy model as an intermediary to efficiently transfer knowledge from a black-box LLM to a smaller model.
Experimental results show that Proxy-KD outperformed existing White-box-based knowledge distillation techniques, significantly improving the efficiency of knowledge transfer from black-box teacher models.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.