
* support image/webp file type Signed-off-by: Elwin <61868295+hzhaoy@users.noreply.github.com> Signed-off-by: Elwin <hzywong@gmail.com> * docs: add webp image format in supported_formats.md Signed-off-by: Elwin <61868295+hzhaoy@users.noreply.github.com> Signed-off-by: Elwin <hzywong@gmail.com> * test: add a test case for `image/webp` file Signed-off-by: Elwin <hzywong@gmail.com> * style: apply styling Signed-off-by: Elwin <hzywong@gmail.com> * test: update test case of converting `image/webp` file with more ocr engines Signed-off-by: Elwin <hzywong@gmail.com> * style: apply styling Signed-off-by: Elwin <hzywong@gmail.com> * rename test file Signed-off-by: Michele Dolfi <dol@zurich.ibm.com> --------- Signed-off-by: Elwin <61868295+hzhaoy@users.noreply.github.com> Signed-off-by: Elwin <hzywong@gmail.com> Signed-off-by: Michele Dolfi <dol@zurich.ibm.com> Co-authored-by: Michele Dolfi <dol@zurich.ibm.com>
36 lines
1.1 KiB
Markdown
Vendored
36 lines
1.1 KiB
Markdown
Vendored
Docling can parse various documents formats into a unified representation (Docling
|
|
Document), which it can export to different formats too — check out
|
|
[Architecture](../concepts/architecture.md) for more details.
|
|
|
|
Below you can find a listing of all supported input and output formats.
|
|
|
|
## Supported input formats
|
|
|
|
| Format | Description |
|
|
|--------|-------------|
|
|
| PDF | |
|
|
| DOCX, XLSX, PPTX | Default formats in MS Office 2007+, based on Office Open XML |
|
|
| Markdown | |
|
|
| AsciiDoc | |
|
|
| HTML, XHTML | |
|
|
| CSV | |
|
|
| PNG, JPEG, TIFF, BMP, WEBP | Image formats |
|
|
|
|
Schema-specific support:
|
|
|
|
| Format | Description |
|
|
|--------|-------------|
|
|
| USPTO XML | XML format followed by [USPTO](https://www.uspto.gov/patents) patents |
|
|
| JATS XML | XML format followed by [JATS](https://jats.nlm.nih.gov/) articles |
|
|
| Docling JSON | JSON-serialized [Docling Document](../concepts/docling_document.md) |
|
|
|
|
## Supported output formats
|
|
|
|
| Format | Description |
|
|
|--------|-------------|
|
|
| HTML | Both image embedding and referencing are supported |
|
|
| Markdown | |
|
|
| JSON | Lossless serialization of Docling Document |
|
|
| Text | Plain text, i.e. without Markdown markers |
|
|
| Doctags | |
|