Link copied to clipboard
Sinhala & Tamil OCR and legacy-font conversion | Sri AI
AI: Document AI & OCR Popular

Sinhala & Tamil OCR and legacy-font conversion | Sri AI

(0 reviews)
Advanced 12 views

Course overview

Training OCR for our own scripts and rescuing decades of documents typed in old non-Unicode fonts such as FM and Shree-Lipi.

Level: Advanced  ·  Mode: Part-time

Who this course is for

Engineers, publishers and archivists working with Sinhala and Tamil documents.

What you will learn

  • Train OCR models on Sinhala and Tamil script
  • Convert legacy-font documents to Unicode
  • Correct OCR output automatically
  • Build a digitisation pipeline for books

Syllabus

  1. Module 1: Script structure and Unicode
  2. Module 2: Legacy fonts and their mappings
  3. Module 3: Training data for local scripts
  4. Module 4: Training and fine-tuning OCR
  5. Module 5: Post-correction with language models
  6. Module 6: Book digitisation case study

Final project

Every module ends in hands-on practice, and the course ends with a project you build and present. Your certificate names that project.

Before you start

OCR fundamentals.

Open-source tools you will use

Rate This Course