AI: Document AI & OCR
Popular
Sinhala & Tamil OCR and legacy-font conversion | Sri AI
Advanced
13 views
Course overview
Training OCR for our own scripts and rescuing decades of documents typed in old non-Unicode fonts such as FM and Shree-Lipi.
Level: Advanced · Mode: Part-time
Who this course is for
Engineers, publishers and archivists working with Sinhala and Tamil documents.
What you will learn
- Train OCR models on Sinhala and Tamil script
- Convert legacy-font documents to Unicode
- Correct OCR output automatically
- Build a digitisation pipeline for books
Syllabus
- Module 1: Script structure and Unicode
- Module 2: Legacy fonts and their mappings
- Module 3: Training data for local scripts
- Module 4: Training and fine-tuning OCR
- Module 5: Post-correction with language models
- Module 6: Book digitisation case study
Final project
Every module ends in hands-on practice, and the course ends with a project you build and present. Your certificate names that project.
Before you start
OCR fundamentals.