AI Briefing
KO

20 Million-Scale Indian Legal Dataset Released

·2026.04.14 23:14

Key point

A large-scale dataset containing over 20 million Indian legal documents, citation graphs, and vector embeddings has been released.

Details

This is a 20+ million legal document dataset including judgments from the Supreme Court of India, 25 High Courts, and 14 tribunals. Built over the past 2 years, it secures structured metadata and citation relationships.

Core Components:

  • Structured Metadata: Includes court, bench, date, parties, judges, cited provisions, related laws, case type, and more.
  • Citation Graph: Provides a machine-readable graph classifying relationships across the entire corpus into 'distinguished', 'overruled', and 'mentioned'.
  • Embeddings and Cross-References: Voyage AI (1024d dense) and BM25 (sparse) vectors have been applied, with cross-references to 23,122 statutes and related case law.

Key Potential Uses:

  • Research and Prediction: Graph neural network (GNN) research, legal outcome prediction, analysis of judgment influence.
  • Language Model Enhancement: Fine-tuning legal-domain-specific Indian language models with formal, precise text rather than news or conversational style.
  • RAG Evaluation: Benchmarking retrieval-augmented generation (RAG) performance in the legal domain using citation relationships as ground truth.

This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.

Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.