Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
From Boilerplate to Boardroom: Unlocking Complex Legal & Financial Docs with PaddlePaddle
Learn how to combine PaddleOCR, layout analysis, and Baidu ERNIE to extract, clean, and structure bilingual legal and financial documents into JSON.
This presentation is a technical deep dive into the engineering choices required to build a robust pipeline for parsing complex legal and financial documents. Using slides with detailed code snippets and comparative outputs, I will walk through our systematic approach.
Layout Analysis: We’ll begin by analyzing the initial output from Baidu’s PP-DocLayout-L on a complex, bilingual document. I’ll then walk through the Python code snippets we engineered to handle specific challenges like multi-column layouts and footnote separation.
Comparative OCR Analysis: I will present a direct, side-by-side comparison of Baidu’s PaddleOCR versus the Tesseract engine on identical legal text. We’ll examine the specific scenarios where one outperforms the other, providing a clear-eyed view of their respective strengths and weaknesses.
LLM-Augmented Structuring: This is where we go beyond standard tools. I will share our tips and tricks for leveraging Large Language Models (LLMs) such as Baidu’s ERNIE to handle the final, most nuanced structuring tasks. You’ll see how we used LLMs to interpret ambiguous contexts and transform unstructured phrases into clean, structured data—a task that pure OCR or regex struggles with.
Final Structuring Pipeline: To conclude, I will showcase the Python script that integrates these components, transforming the processed text into a final, hierarchically correct JSON object, ready for analysis or downstream tasks.
Sigtica offers containerized data engineering, ML, and database development tools.
Compose Email
Loading recent emails...