AI Briefing
KO

Medical NLP Stack OpenMed 2.0 Released

·2026.07.29 07:08

Key point

OpenMed 2.0, an open-source NLP stack for protecting medical data, has been released with multilingual PHI de-identification capabilities.

Details

OpenMed 2.0, a medical NLP stack, has been released under the Apache-2.0 license. The core of this update is PHI (Protected Health Information) de-identification functionality that can be performed in a local environment without sending data to cloud APIs.

Key features include:

  • Lightweight models: Uses encoders with 279M and 560M parameters to perform efficient de-identification tasks, offering higher security than cloud-based APIs.
  • Expanded multilingual support: Language coverage has expanded from 23 to 56, with newly added PII models for Bengali (279M), Chinese (560M), and Tamil (279M).
  • Browser execution support: Provides artifacts that can run directly in browser environments via Transformers.js.
  • Identifier validation: Includes validation for 51 National Identifiers, supporting sophisticated validation that uses actual checksums rather than simple regular expressions.
  • Data anonymization risk management: Beyond simple identifier removal, it introduces a structured release risk workflow that assesses re-identification risk through data linkage and helps prevent it.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.