AI Briefing
KO

Vision capability introduced to the fine-tuning API

·2024.10.01 19:04

Key point

OpenAI has launched vision fine-tuning, a feature that lets developers train GPT-4o using both images and text together.

Details

Vision fine-tuning has been introduced for GPT-4o, allowing developers to now customize the model's training using datasets that include images in addition to text. This makes it possible to maximize performance in fields where visual understanding is essential, such as visual search, autonomous driving object detection, and medical imaging analysis.

The training method is similar to existing text fine-tuning, and is done by uploading an appropriately formatted image dataset. Performance improvements are possible with as few as 100 images, and more refined results can be achieved as the amount of data increases.

Key use cases are as follows:

  • Grab: By training on road images, they improved lane count accuracy by 20% and speed limit sign location detection accuracy by 13%.
  • Automat: By leveraging screenshot data, they raised the success rate of an RPA agent that recognizes UI elements from 16.60% to 61.67%.
  • Coframe: By training on images and code together, they improved the match rate of visual style and layout when generating websites by 26%.

OpenAI continuously performs automated safety evaluations on fine-tuned models, and guarantees data security through its Enterprise privacy commitment.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.