Packages

Elixir bindings for anydoc: convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF documents into clean GitHub-Flavored Markdown.

Current section

Files

Jump to
anydoc_ex README.md
Raw

README.md

# anydoc_ex
[![CI](https://github.com/kentaro/anydoc_ex/actions/workflows/ci.yml/badge.svg)](https://github.com/kentaro/anydoc_ex/actions/workflows/ci.yml)
[![Hex.pm](https://img.shields.io/hexpm/v/anydoc_ex.svg)](https://hex.pm/packages/anydoc_ex)
[![Docs](https://img.shields.io/badge/hex-docs-blue.svg)](https://hexdocs.pm/anydoc_ex)
[![License: MIT](https://img.shields.io/badge/license-MIT-blue.svg)](LICENSE)
Convert Word, PowerPoint, Excel, OpenDocument, RTF, EPUB, CSV, and PDF files into clean GitHub-Flavored Markdown. Elixir bindings for the [anydoc](https://github.com/firecrawl/anydoc) Rust crate, built by [Firecrawl](https://firecrawl.dev).
Every format parses into one shared document model and renders through a single Markdown serializer, so headings, tables, lists, and footnotes come out the same no matter which format goes in. Conversion runs as a Rustler NIF on dirty CPU schedulers, so it never blocks the BEAM's normal schedulers — median conversion time upstream is under 5 ms per document.
## Installation
Add `anydoc_ex` to your dependencies:
```elixir
def deps do
[
{:anydoc_ex, "~> 0.1"}
]
end
```
The native extension compiles from source at build time, so a [Rust toolchain](https://rustup.rs) is required. Precompiled binaries ([rustler_precompiled](https://hex.pm/packages/rustler_precompiled)) are on the roadmap.
## Usage
```elixir
# From a file path:
{:ok, markdown} = Anydoc.to_markdown("report.docx")
# From bytes, with the format detected from the content:
{:ok, markdown} = Anydoc.to_markdown_bytes(bytes)
# Or name it, which signature-less formats (CSV) need:
{:ok, markdown} = Anydoc.to_markdown_bytes(bytes, :csv)
# Or stop at the document model, which also carries embedded assets:
{:ok, %Anydoc.Document{markdown: markdown, assets: assets}} = Anydoc.to_document(bytes)
```
Each function has a bang variant (`Anydoc.to_markdown!/1`, `Anydoc.to_markdown_bytes!/2`, `Anydoc.to_document!/2`) that returns the value directly and raises `Anydoc.Error` on failure.
## Supported formats
| Format | Atom | Extensions |
| ---------------- | ---------------------- | ---------------------------------------------------------- |
| Word | `:doc`, `:docx` | `.doc`, `.docx`, `.docm` |
| PowerPoint | `:ppt`, `:pptx` | `.ppt`, `.pps`, `.pot`, `.pptx`, `.pptm`, `.ppsx`, `.ppsm` |
| Excel | `:excel` | `.xls`, `.xlsx`, `.xlsm`, `.xlsb` |
| OpenDocument | `:odt`, `:ods`, `:odp` | `.odt`, `.ods`, `.odp` |
| Rich Text Format | `:rtf` | `.rtf` |
| EPUB | `:epub` | `.epub` |
| CSV | `:csv` | `.csv` |
| PDF | `:pdf` | `.pdf` |
Container variants that share a parser map onto one atom: `.docm` is `:docx`, `.pptm`/`.ppsx`/`.ppsm` are `:pptx`, and every Excel container is `:excel`.
## Errors
A conversion fails only when no meaningful Markdown could come out of the file. The error's `:code` names what went wrong:
```elixir
case Anydoc.to_markdown(path) do
{:ok, markdown} ->
markdown
{:error, %Anydoc.Error{code: code}} when code in [:encrypted, :unsupported] ->
# No document comes out of these, so record the file and take the next one.
nil
{:error, error} ->
raise error
end
```
| Code | Meaning |
| ----------------- | ------------------------------------------------------------------- |
| `:unsupported` | Unknown format, or one that cannot be converted (an image-only PDF) |
| `:malformed` | Structurally unusable: no meaningful content could be extracted |
| `:encrypted` | Encrypted or password-protected |
| `:resource_limit` | Crossed a fixed safety limit (decompression, nesting, node count) |
| `:missing_part` | A part required for any meaningful output is absent |
| `:io` | The file could not be read, from `Anydoc.to_markdown/1` only |
`Anydoc.Error` is an exception struct, so the same value works in pattern matches and in `raise`. `:part` names the package part at fault (`:malformed`, `:missing_part`), `:limit` names the safety limit crossed (`:resource_limit`), and `Exception.message/1` carries the whole message. Passing anything but a known format atom raises `FunctionClauseError`.
## Format detection
The format is read from the file content, using the marker its specification designates: the PDF header, the RTF open group, OLE stream names, the ZIP package mimetype and content types. CSV has no such marker, so detection returns `nil` for it and the extension, or an explicit format, names it instead.
```elixir
Anydoc.format_from_bytes(bytes) # :docx, or nil when nothing matches
Anydoc.format_from_extension(".pptm") # :pptx
Anydoc.format_from_path("report.odt") # :odt
```
## Images and embedded objects
Markdown cannot embed bytes, so an embedded image renders as its alt text while the bytes stay on `Anydoc.Document.assets`, tagged with a media type and the package part they came from. Images that carry an external URL render as ordinary Markdown images.
```elixir
{:ok, document} = Anydoc.to_document(File.read!("report.docx"))
for %Anydoc.Asset{id: id, media_type: media_type, bytes: bytes} <- document.assets do
File.write!("asset-#{id}#{extension_for(media_type)}", bytes)
end
```
The upstream crate exposes a full block-level document model (headings, tables, lists, notes); this binding does not mirror that tree yet — `Anydoc.Document` carries the rendered Markdown plus the embedded assets. The full model is a candidate for a future release.
## Development
```bash
mix deps.get
mix test # compiles the Rust NIF on first run
mix format --check-formatted
mix credo --strict
mix dialyzer
```
The NIF crate lives in `native/anydoc_nif` and wraps only the public API of the `anydoc` crate. Test fixtures under `test/fixtures` are self-describing sample documents; `with_image.docx` is a minimal hand-built OOXML package with one embedded PNG, used to exercise the asset path.
Full behavior notes and benchmarks live in the [anydoc repository README](https://github.com/firecrawl/anydoc#readme).
## License
[MIT](LICENSE)