Google Unveils Open VLM PaliGemma
Key point
Google has released PaliGemma, a new open vision language model (VLM) that combines SigLIP and Gemma.
Details
Google has released PaliGemma, a new vision language model (VLM) family that combines a SigLIP-So400m image encoder with a Gemma-2B text decoder.
The model is provided in three checkpoints depending on use case:
- PT (Pretrained): a pretrained model for fine-tuning on specific tasks
- Mix: a model mixed across various tasks for general inference and research
- FT (Fine-tuned): fine-tuned models specialized for specific academic benchmarks
Users can choose from three resolutions—224x224, 448x448, 896x896—and various precisions (bfloat16, float16, float32). Higher-resolution models are advantageous for fine-grained tasks such as OCR but require more memory.
Rather than being a conversational model, PaliGemma is designed to deliver optimal performance when fine-tuned for specific tasks such as image captioning, visual question answering (VQA), object detection, and segmentation.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.