![]() * Keep page.parsed_page.textline_cells and page.cells in sync, including OCR Signed-off-by: Christoph Auer <cau@zurich.ibm.com> * Make page.parsed_page the only source of truth for text cells Signed-off-by: Christoph Auer <cau@zurich.ibm.com> * Small fix Signed-off-by: Christoph Auer <cau@zurich.ibm.com> * Correctly compute PDF boxes from pymupdf Signed-off-by: Christoph Auer <cau@zurich.ibm.com> * Use different OCR engine order Signed-off-by: Christoph Auer <cau@zurich.ibm.com> * Add type hints and fix mypy Signed-off-by: Christoph Auer <cau@zurich.ibm.com> * One more test fix Signed-off-by: Christoph Auer <cau@zurich.ibm.com> * Remove with pypdfium2_lock from caller sites Signed-off-by: Christoph Auer <cau@zurich.ibm.com> * Fix typing Signed-off-by: Christoph Auer <cau@zurich.ibm.com> --------- Signed-off-by: Christoph Auer <cau@zurich.ibm.com> |
||
---|---|---|
.. | ||
docx | ||
json | ||
xml | ||
__init__.py | ||
abstract_backend.py | ||
asciidoc_backend.py | ||
csv_backend.py | ||
docling_parse_backend.py | ||
docling_parse_v2_backend.py | ||
docling_parse_v4_backend.py | ||
html_backend.py | ||
md_backend.py | ||
msexcel_backend.py | ||
mspowerpoint_backend.py | ||
msword_backend.py | ||
pdf_backend.py | ||
pypdfium2_backend.py |