AI Briefing
KO

On-device AI Face Verification Pipeline Optimization

·2026.01.23 09:00

Key point

On-device Android face verification optimization improved latency by 37% and throughput by 530%.

1 / 2

Details

Hyperconnect's Match Group AI team developed on-device photo recommendation technology to find profile-suitable images among thousands of photos in a gallery. For privacy protection, all processing had to be completed within the device, and the speed of the Face Verification pipeline, the first step in this process, became a key challenge.

Measurements were conducted on a Galaxy S24 (Android 14) with TensorFlow Lite CPU acceleration. Noise was reduced through 500 repeated measurements and device cooling. In the initial results, Load Image took 46.4ms, Face Detection took 30.5ms, and 3rd-party Face Recognition took 61.3ms, for a total of 138.2ms, with throughput at 7.2 photos/s.

The first area addressed was Face Detection pre/post-processing. Previously, all candidate bounding boxes were decoded first before discarding low-confidence candidates, but this was changed to filter-first, greatly reducing the number of candidates to decode. Additionally, a fused operation combining sorting and Top-K selection was applied, reducing decoding time from 8.9ms to 1.9ms, improving overall latency from 76ms to 70ms, and throughput from 13 photos/s to 14 photos/s.

The next bottleneck was takePromisingTopKIndices. The existing filter -> sortedByDescending -> take approach caused full sorting and intermediate collection creation. Switching this to a Min Heap-based algorithm that only maintains the top K elements improved the response time of the Pre NMS Filter & TopK stage by 78%.

Another bottleneck was the TensorBuffer.getFloatValue call. Since ByteBuffer.getFloat incurred high repeated call costs due to going through JNI, this was switched to using TensorBuffer.floatArray. As a result, Pre NMS Filter & TopK, Decode Bboxes, Decode Landmarks, and Apply NMS each became about 10-20% faster, and the overall pipeline latency decreased from 70.2ms to 69ms.

After optimizations outside the model, the TensorFlow Lite thread pool size was adjusted. 1 thread was faster than the default 4 threads, and inference P50 actually improved from 13.8ms to 6.2ms. Profiling results showed that some operations actually became 7.7-7.8x slower as thread count increased, confirming that parallelization is not always beneficial.

Finally, throughput was boosted through model instance parallelization. The face detection model was processed in parallel by increasing the number of interpreters, and the face recognition model was also parallelized with a separate thread pool. As a result, the overall pipeline throughput increased from 16.6 photos/s to 37.9 photos/s to 45.9 photos/s. Ultimately, per-image latency decreased from 138ms to 87ms, approximately 37% reduction, and throughput improved from 7.2 photos/s to 45.9 photos/s, approximately 530% improvement.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.