Current section
Files
Jump to
Current section
Files
erlang_python
CHANGELOG.md
CHANGELOG.md
# Changelog
## 5.0.1 (2026-09-22)
### Fixed
- Tasks on the event loop pool that awaited `asyncio.sleep` for more than a
few milliseconds never completed when several ran at once: with 24
concurrent 50 ms sleeps, 23 timed out, while the same tasks run one after
another always finished (in Hornbeam, concurrent ASGI requests hung until
the request timeout). Every pool loop scheduled its timers on the default
loop and polled that loop's queue, where the per-loop callback ids
collided. Each pool loop now drives a Python loop bound to its own
resource, a timer expiry is dispatched to the loop that set it, and a loop
with no tasks leaves its pending events to whoever polls it.
- A Python function called with `py:call` that called `erlang.call`, where
the Erlang callback called `py:call` again, hung until the request timeout
and then failed with "callback synchronisation lost; retry" or
`{error, timeout}`, even one level deep; the same chain started with
`py:eval` worked. `erlang.call` blocked the context thread on the thread
worker pipe and the nested `py:call` waited on that context. The context
thread now serves its own requests while it waits for the callback, so the
nested `py:call` runs inline on the same context, at any depth and with
several sequential `erlang.call`s in one function; the Python code still
runs exactly once. Inside a running asyncio loop (`py_context:start_loop/2`,
`asyncio.run` in a called function) the thread path stays.
- The shared buffer's flow control and the sync `erlang.sleep` call Erlang
through a blocking path that never suspends the request, so a `py:eval`
that reads a shared buffer or sleeps is not replayed around them.
- After a callback resumed, a `py:eval` or `py:call` that used a function
defined with `py:exec` failed with `NameError: name ... is not defined`:
the replay ran in the context globals instead of the caller's namespace.
The replay now runs in the namespace of the original request.
- An Erlang callback could not see the caller's `__main__` functions and
was routed to another context: `py:call('__main__', double, [X])` from a
callback failed with "module '__main__' has no attribute 'double'". The
callback process is now bound to the suspended context and shares the
caller's namespace, so the README reentrant example passes as written.
## 5.0.0 (2026-08-29)
### Added
- **`isolated` context mode** - `py_context:new(#{mode => isolated})` runs
CPython in a child OS process per context, with the same `call/eval/exec`,
callback, `erlang.send`/`whereis`, worker-loop and pool API as the embedded
modes. It is the first mode with a hard bound: `py_context:interrupt/1`
stops a blocking C call (a signal in the child) and `SIGKILL` is the
backstop after `kill_after` ms; `py_context:kill/1` kills at once. `rlimits`
(`as`, `cpu`, `nofile`) and a cgroup v2 directory bound the child; a
segfault in a C extension returns `{error, {child_exited, {signal, 11}}}`
and the node survives. The child restarts on crash within a budget
(`restart`, `max_restarts`, `restart_period`); `py_context:child_info/1`
reports its OS pid. Children are reaped by the VM and exit when the BEAM
dies (socket EOF watchdog, `PR_SET_PDEATHSIG` on Linux, `PROC_PDEATHSIG_CTL`
on FreeBSD). `cgroup` is refused outside Linux; rlimits apply everywhere:
`as` is kernel-enforced on Linux and FreeBSD and enforced by an RSS
watchdog in the child on macOS (`{child_exited, {memory_limit, Bytes}}`).
Validated on macOS (arm64) and FreeBSD 14.3 (OTP 28, Python 3.11).
- **`py_context:pass_fd/2`** - hands a file descriptor to an isolated child
over the control socket (`SCM_RIGHTS`), so `erlang.server.serve` works out
of process: Erlang binds once, N killable children accept.
- **Pure-Python ETF codec** (`priv/_erlang_impl/_etf.py`) with the type
mapping of `py_convert.c`; the child needs no C extension. Integers beyond
64 bits round-trip exactly in isolated mode.
- **Shared memory** - `py_shm:new/1,2`, `write/3`, `read/3`, `binary/3`
(no copy), `close/1`: fixed-size regions over
[iommap](https://hex.pm/packages/iommap) (optional dependency) that any
context mode maps as `erlang.SharedMemory` (buffer protocol, numpy
friendly). `py_buffer:new(#{shared => true})` is a streaming buffer over
such a region with ring backpressure, usable as `wsgi.input` in isolated
contexts. Handles are plain terms and travel inside any argument or result;
`py_shm:read_only/1` and `new(Size, #{writable => false})` hand Python a
read-only mapping. `py_buffer:write/3` takes a timeout for the case where
the ring is full and nobody reads (default 30 s).
- `py:python_executable/0`, `py:kill/1`, `py_nif:os_kill/2`.
- `py_isolated` is a `gen_statem` (states `idle`, `{busy, Id}`, `looping`,
`stopping_loop`, `{restarting, Reason}`): `sys:get_state/1` and
`sys:trace/2` work on isolated contexts, requests arriving during a
restart are served by the new child, and `py_context:kill/1` returns once
the new child is up.
- Timeouts on an isolated context cancel their own request only (queued
requests are dropped, the executing one is interrupted); the kill backstop
is bound to that request, so a busy shared context is never killed because
another caller gave up. Soak-tested: callback storms, interrupt/kill
storms, loop churn, 60 s mixed workload with resource counters checked.
- Guide: `docs/isolated.md`, with what each of the three modes guarantees.
### Changed
- The NIF side of every context request goes through one dispatcher
(`ctx_dispatch`, `ctx_dispatch_async` in `c_src/py_nif.c`) instead of a
per-request copy of the enqueue-and-wait loop; the execute functions are
`ctx_execute_*` and the thread functions `ctx_thread_main_*`, since both
serve worker and owngil contexts. Creating a process-local env and
applying imports or paths run on the context thread in `worker` mode too;
the scheduler-side copies of those paths are gone.
- The NIF function table is assembled from one `PY_*_NIFS` macro per area,
defined at the end of the file that owns the NIFs.
- `py_context` keeps the API and the reply protocol; the process body for
embedded modes moved to `py_context_embedded`. `py` delegates streaming,
virtual environments and shared dicts to `py_stream`, `py_venv` and
`py_shared_dict`. The public API is unchanged.
### Documentation
- `make check-code-map` (also run by CI) verifies that every source file is
in `docs/code-map.md`, every Erlang module has a moduledoc and a row in
the Modules table of `test/coverage_audit.md`.
### Removed
- The legacy worker API (`py_nif:worker_new/0,1`, `worker_call`, `worker_eval`,
`worker_exec`, `worker_next`, `worker_destroy`, `import_module/2`,
`get_attr/3`, `set_callback_handler/2`, `send_callback_response/2`,
`resume_callback/2`) and the single executor thread behind it, the
`async_worker_*`/`async_call`/`async_gather`/`async_stream` NIFs that only
returned `deprecated`, the unused worker pool (`pool_*` NIFs), the
`cancel_reader/writer` aliases, and the unreachable inline executor
branches of the context NIFs. Contexts (`py_context`, `py:call/3`) are the
only execution path. `py:memory_stats/0` and `py:gc/0,1` now run on the
calling scheduler under the GIL.
### Fixed
- `pthread_timedjoin_np` was called without `_GNU_SOURCE`, an implicit
declaration on Linux that newer compilers reject.
- Callback pipes waited with `select()`, which is undefined for a file
descriptor above 1024: in a VM with many open files a thread callback
could time out with "Failed to spawn thread handler". The waits use
`poll()`, the handler ready-wait no longer holds the GIL, and
`py_thread_handler` logs a failed ready signal.
## 4.1.0 (2026-08-15)
### Added
- **Worker loops** - `py_context:start_loop/1,2` runs an `ErlangEventLoop`
forever on the context thread and returns at once; `py_context:submit/4,5`
and `submit_await/4,5,6` schedule a coroutine or function on it from Erlang
(results as `{async_result, TaskRef, _}`, `py_event_loop:await/1,2`);
`stop_loop/1,2` stops it cooperatively, then interrupts after a grace
period; `loop_ref/1` exposes the loop for `py_nif:submit_task/7`. The owner
receives `{py_loop_exit, Ctx, Result}` when the loop ends. While a loop
runs, `call/eval/exec/call_method` on that context return
`{error, loop_running}`. `py_context:new/1` takes `preload => Code`, run once
in the context before anything else. See `docs/workers.md`.
- **`erlang.server`** - `serve(listen_fd, protocol_factory, udp=False)`,
`adopt(fd, protocol_factory)` and `stop_serving(server)`: serve TCP or UDP
on a socket Erlang bound (`py:dup_fd/1` per worker) or take over one
accepted connection, from a coroutine scheduled with `submit`. This is the
gunicorn shape inside the VM: Erlang binds once, N owngil contexts accept
on their copy of the fd, Erlang supervises, scales and reloads.
- **Injection into subinterpreter loops** - `py_nif:process_ready_tasks/1`
attaches a thread state to the loop's subinterpreter, so `submit_task` works
for owngil loops, idle or running; scheduling into a running loop now wakes
it instead of waiting for the next poll timeout (about 25 us round trip
instead of up to 1 s). Tasks that fail to start (missing module or function,
argument conversion, the call itself raising) are reported to the caller as
`{async_result, Ref, {error, Reason}}` instead of being dropped.
### Fixed
- **owngil contexts had no event loop** - `owngil_context_thread_main` created
the `py_event_loop` module without a default loop, so `erlang.run()`,
`create_server` and channels raised "Erlang event loop not initialized" in
owngil contexts. Each owngil context now owns an `ErlangEventLoop` served by
its own `py_event_worker`. `py_nif:context_get_event_loop/1` returned the
main interpreter's loop for owngil contexts, which made every owngil start
re-point the main loop's worker to a process that died with the context.
- **owngil dispatch** - calls into owngil contexts went through a blocking
dispatch on a dirty CPU scheduler with a 30 s cap
(`OWNGIL_DISPATCH_TIMEOUT_SECS`); they now use the same async queue as
worker mode: no dirty scheduler held during the call, no cap, and a lower
round trip (about 11.6 us against 15.6 us before on the bench machine).
- **fd closed while still in the poll set** - transports closed their socket
right after `ERL_NIF_SELECT_STOP` was issued, which under connection churn
produced `Bad input fd in erts_poll()` and `enif_select ... stealing
control of fd` reports and could deliver events to the wrong resource once
the number was reused. Transports now detach the fd and hand it to the NIF
(`_release_fd_resource(fd_key, take_ownership)`), which closes it from the
select stop callback; the reselect path and the close path serialise on the
loop mutex. 10k connections across four workers now log nothing.
- **Queued tasks dropped in pairs** - `py_nif:process_ready_tasks/1` dequeued
one whole iovec element per task; when erts stored several small task
binaries in one element (tasks queued behind a busy worker), every task
after the first in that element was lost, seen as every second
`submit_task` never answering on slow machines. It now dequeues exactly the
bytes each term consumed. Tasks beyond the batch limit queued behind a
running loop were also left waiting for the next wakeup; the running-loop
path now returns `more` like the idle path.
- **Recycled handles cancelled by their previous owner** - `ErlangEventLoop`
handed pooled `Handle` objects out of `call_soon` (and out of `call_at`
when the delay rounded to zero, which `asyncio.sleep(0.001)` does depending
on the clock value). asyncio cancels such handles after they ran
(`sleep` does in its `finally`), which cancelled whatever callback had been
given the recycled handle since: every second sleeper never woke. Only the
fd event handles created inside `_dispatch` are pooled now; `call_soon`
and `call_at` return fresh handles.
- **Re-arming a read select from the Python thread** - re-selecting READ on
an fd the BEAM had moved into a scheduler poll set crashed inside
`enif_select` when done from the loop thread (transport `resume_reading`,
`add_reader` on an fd with an active writer). Read re-arms now go through
the loop's `py_event_worker` (`py_nif:fd_arm/2`).
## 4.0.0 (2026-08-15)
### Breaking Changes
- **Callback results returning an Erlang string** - a string returned from an
Erlang callback (`"abc"`, a list of integers) now reaches Python as
`[97, 98, 99]` rather than `'abc'`, the same conversion call arguments have
always used. Return a binary (`<<"abc">>`) for a Python `str`. Nothing raises
on upgrade, so check what your registered callbacks return before deploying.
See the encoding change under Changed below.
### Added
- **Interrupt running Python** - `py:interrupt/1` and `py_context:interrupt/1`
raise `KeyboardInterrupt` in the thread executing a context; the in-flight call
returns `{error, interrupted}`. `py_context:call/eval/exec` now interrupt
automatically when their timeout expires, so `{error, timeout}` stops the Python
code instead of only abandoning the reply while the thread kept burning CPU.
Works in both `worker` and `owngil` modes, and is callable while the context
process is blocked in a NIF. CPython delivers async exceptions at bytecode
boundaries, so code blocked in a C call (`time.sleep`, a numpy kernel, a socket
read) is interrupted once that call returns. See `docs/interrupts.md`.
- **Per-context memory caps** - `py_context:new(#{mode => owngil, memory_limit => Bytes})`
caps memory allocated by one context; exceeding it raises `MemoryError` there and
leaves other contexts untouched. Requires `{enable_memory_limits, true}` in the
application environment, since the allocator is hooked before Python starts.
Accounting covers obmalloc traffic only: allocations over 512 bytes (large
binaries, numpy buffers) are not counted, granularity is one 1 MB arena, and
`worker` mode returns `{error, memory_limit_requires_owngil}` because those
contexts share the main interpreter. `py_nif:context_memory_usage/1` reports
usage. See `docs/memory.md`.
- **Async generator streaming** - `py:stream_start/3,4` accepts async generators
again, driving them on a private event loop and emitting the same
`{py_stream, Ref, ...}` events. The docs claimed this since 3.0.0 without the
code behind it. `py:stream/4` with kwargs and `py:stream_eval/1,2` remain
sync-only, which is now stated explicitly.
### Changed
- **Callback results cross as external term format** - results returned from an
Erlang callback into Python are encoded with `term_to_binary` and decoded by the
same `term_to_py` converter used for call arguments, replacing the Python-repr
string that was parsed with `ast.literal_eval`. This fixes binaries containing
backslashes, quotes, newlines or tabs (which produced an unparseable literal and
were silently handed to Python as the raw repr text), `[]` arriving as `''`,
float precision loss, and the base64 round-trip for pids and references, which
now cross as native `Pid` and `Ref` objects.
Breaking: an Erlang string returned from a callback (`"abc"`, a list of
integers) now reaches Python as `[97, 98, 99]` rather than `'abc'`, the same
conversion call arguments have always used. Return a binary (`<<"abc">>`) for a
Python `str`.
## 3.1.1 (2026-05-31)
### Changed
- **Lower minimum OTP to 27** - `minimum_otp_vsn` is now `27`. The OTP 28/29
support work was source-compatible with 27 (the `try ... catch` cleanups build
fine there), so the floor was raised further than needed. CI now also builds and
runs the full Common Test suite on OTP 27 across Python 3.12/3.13/3.14.
## 3.1.0 (2026-05-30)
### Fixed
- **NIF robustness hardening** - `make_py_error` no longer passes a NULL message/type
to `enif_make_string`/`enif_make_atom` when a Python exception's text isn't
UTF-8-encodable; `binary_to_string` rejects names/code containing an embedded NUL
(which would silently truncate a module/function/attr/code string) rather than
truncating; a leaked `split` method object in the reactor buffer is released; and a
stray debug `fprintf` on the normal worker send path is removed.
### Security
- **No shell for venv/installer commands** - `py:ensure_venv` and dependency
installation now run the executables via `open_port({spawn_executable, ...})` with an
argument list instead of building a shell string for `os:cmd`. Venv paths, requirement
files, and extras are passed literally, so shell metacharacters can't be injected. For
`uv`, `VIRTUAL_ENV` is passed via the port `{env, ...}` option rather than a shell prefix.
- **Bounded shared state + safe stream/log builders** - `py_state` gained an optional
`max_state_entries` cap (default `infinity`, unchanged behavior) enforced with atomic
admission so Python-driven `state_set` can't exhaust node memory, and its size counter
is protected from corruption. The `py:stream` and logging helpers that build Python
source now strictly validate module/function/kwarg names as identifiers (rejecting
injection at positions where quoting is meaningless) and escape string-literal values
including control characters.
- **Validated event-loop fd handles** - The asyncio reader/writer integration no longer
hands Python a raw `fd_resource` pointer as an integer key. Each handle is an opaque id
validated against a registry on every use, so a stale, duplicate, or fabricated id is a
safe no-op (or clean error) instead of a double-free or arbitrary-pointer dereference
that crashed the node. `fd_read`/`fd_write` also moved to dirty IO schedulers.
- **OWN_GIL worker robustness** (Python 3.14+) - A per-request allocation failure in
a subinterpreter worker no longer `break`s (and permanently kills) the worker command
loop; it returns an error and keeps serving. The `owngil_*` dispatch NIFs now run on
dirty IO schedulers and use non-blocking, deadline-bounded pipe reads and writes, so a
stalled or dead worker can't wedge a scheduler forever. The internal `SuspensionRequired`
exception is now looked up per-interpreter (like `ProcessError`), avoiding cross-
interpreter object use under OWN_GIL.
- **Callback suspend/resume lifetime hardening** - The worker resource is now kept
alive for the lifetime of a suspended callback (it could previously be GC'd mid-
suspension, causing a use-after-free on resume). A resume frees any prior result
before storing a new one (no leak/double-replay on a duplicate resume), the
pending-callback thread-local is cleared at the worker request boundary, and the
callback-response pipe writes run on dirty schedulers with non-blocking, deadline-
bounded writes so a stalled reader or large payload can't wedge a scheduler or
desync the framed protocol.
- **Zero-copy buffer pinning** - `py_buffer` no longer relocates (and frees) its
storage while a Python `memoryview` points into it. A write that would grow the
buffer while a view is held now returns an error instead of dangling the view into
freed memory (a use-after-free that crashed the whole node).
- **Bounded recursion in type conversion** - The Erlang<->Python converters now cap
nesting depth, so a deeply nested term (or Python structure) returns a clean error
instead of overflowing the C stack and crashing the whole node.
- **NULL-checked tuple allocation** - Argument-tuple allocations in the call/eval paths
are checked before use, and the Python->Erlang map conversion is bounded against
mid-iteration dict mutation, closing two ways an allocation failure or re-entrant
`__str__` could corrupt memory.
- **Safe term decoding at the NIF boundary** - All `enif_binary_to_term` calls now
pass `ERL_NIF_BIN2TERM_SAFE`, preventing attacker-influenced data (notably a Python
`"__etf__:<base64>"` callback result) from minting new, non-GC'd atoms and exhausting
the atom table. Local-node pids/refs and already-existing atoms still round-trip
unchanged; only brand-new atoms, remote-node pids/refs, and external funs in
Python-supplied payloads are now rejected.
### Changed
- **Support Erlang/OTP 28 and 29** - Validated builds and the full Common Test
suite on OTP 28 and 29. Minimum supported OTP is now 28 (`minimum_otp_vsn`).
CI tests OTP 28 and 29 across Python 3.12/3.13/3.14.
- Replaced deprecated `catch Expr` cleanup calls with `try ... catch ... end`
to silence the new OTP 29 default warning; behavior is unchanged.
## 3.0.0 (2026-05-03)
### Breaking Changes
- **Simplified execution model** - Only two public execution modes: `worker` and `owngil`
- `worker`: Dedicated pthread per context with stable thread affinity (default)
- `owngil`: Dedicated pthread + subinterpreter with own GIL (Python 3.14+)
- Removed `multi_executor` and `free_threaded` from public API
- Internal capability detection still tracks Python features
- **Removed `py:num_executors/0`** - Contexts now use per-context worker threads
instead of a shared executor pool. This function is no longer needed.
- **`py:execution_mode/0` returns `worker | owngil`** - Based on the `context_mode`
application configuration. Previously returned internal capabilities like
`free_threaded`, `subinterp`, or `multi_executor`.
- **Removed `py:async_stream/3,4`** - Streaming async generators was never
implemented behind the API and always returned `{error, stream_not_implemented}`.
Use `py:stream_start/3,4` for sync generators; async-generator support may
return in a later release.
- **Removed `num_executors` / `num_async_workers` configuration** - Both keys
were no-ops after the v3.0 worker rework. Configure context count via
`num_contexts` and the rate-limit ceiling via `max_concurrent`.
- **Strict context-mode validation at the NIF boundary** - `py_nif:context_create/1`
now returns `{error, {invalid_mode, Atom}}` for anything other than `worker | owngil`.
Previously, callers that bypassed `py_context` (notably `py_reactor_context`)
silently mapped any unknown atom — including legacy `auto` and `subinterp` —
to worker mode. Code that relied on that loophole must pass `worker` (or
`owngil`) explicitly.
### Fixed
- **`py:async_call/3,4` + `py:async_await/1,2` round-trip** - Previously the
await receive matched `{py_response, _, _}` while the event loop sent
`{async_result, _, _}`, causing every async call to silently time out.
Async calls now go directly through `py_event_loop:create_task` and
`py_event_loop:await`.
- **`py:async_gather/1,2` actually executes** - Reimplemented as concurrent
`async_call` submission with sequential `async_await`. Returns
`{ok, [Result1, ...]}` on success or `{error, {gather_failed, [{Idx, Reason}, ...]}}`
if any call fails. The previous implementation returned `gather_not_implemented`.
- **Thread-callback flakes (issue #63)** - Six layered defects in the
`erlang.call`/`erlang.async_call` plumbing could deliver wrong values to
the wrong caller under load. Reads now loop on partial/EINTR with a
monotonic deadline; sync writes use a single length-prefixed frame on a
dirty I/O scheduler with deadlined non-blocking writes; the sync wire
carries the originating callback id and the receiver discards mismatched
frames; the async pipe has one writer process per fd with an
atomics-bounded mailbox (`?ASYNC_WRITER_MAX_QUEUE = 10000`) and a
resumable nonblocking parser on the read end; workers that fail to
resync are unlinked from the pool, freed, and bounded by
`MAX_POISONED_WORKERS = 64`.
### Documentation
- Audited every fenced code block in `README.md` and `docs/*.md` for
current-API references. Fixed `Py_GIL_OWN` to `PyInterpreterConfig_OWN_GIL`
in `docs/scalability.md`, corrected the `multi_executor` fallback claim
in `docs/migration.md`, and repaired a broken `SharedDict` example in
`docs/shared-dict.md`.
- New `test/coverage_audit.md` maps every public `py:*` and `erlang.*` API
to its test suite. Added cases for `py:cast/4`, `py:async_gather/2`, and
`py:dup_fd/1` so each documented API has a regression test.
- New `scripts/lint_doc_snippets.escript` (driven by `make lint-docs` and
CI) statically validates every Erlang `py:Fn(/N)` call and parses every
Python block in the docs. Snippets that intentionally show removed APIs
or REPL output opt out via `<!-- skip-lint -->`.
### Changed
- **Per-context worker threads** - Each context now gets its own dedicated pthread
that handles all Python operations. This provides stable thread affinity for
numpy/torch/tensorflow compatibility without needing a shared executor pool.
- **Async NIF dispatch** - Context operations use async NIFs with message passing
instead of blocking dirty schedulers. This improves concurrency under load.
- **Request queue per context** - Replaced single-slot request pattern with proper
request queues that support multiple concurrent callers.
- **No global asyncio policy install on Python 3.14+.** `asyncio.set_event_loop_policy`
was deprecated in 3.14 and is removed in 3.16. The Erlang integration's run path
already uses `loop_factory=` (`erlang.run/1`, `asyncio.Runner`) so the global
policy was only a convenience for bare `asyncio.run()` inside `py:exec`. We now
skip the install on 3.14+ to avoid the deprecation warning. On 3.14+ use
`erlang.run(main)` or `asyncio.Runner(loop_factory=erlang.new_event_loop)`
explicitly. Behavior on Python 3.9–3.13 is unchanged. `erlang.install()` raises
`RuntimeError` on 3.14+ (still emits a `DeprecationWarning` and works on 3.12–3.13).
### Removed
- Multi-executor pool (`g_executors[]`, `multi_executor_start/stop`)
- `context_dispatch_call/eval/exec` functions (dead code)
- References to `PY_MODE_MULTI_EXECUTOR` in context operations
- `py_async_pool` legacy gen_server (unused after async API rewire)
- `priv/_erlang_impl/_ssl.py` (`SSLTransport`, `create_ssl_transport`) had no
importer and was never wired into the asyncio event loop. Removed.
- Internal `py_util` exports `send_response/3`, `normalize_timeout/1`, and
`normalize_timeout/2` had no callers anywhere. Removed. The module is
marked `@private`; no external API changes.
- **Explicit `py:subinterp_*` handle API removed.** `py:subinterp_create/0`,
`subinterp_destroy/1`, `subinterp_call/4,5`, `subinterp_eval/2,3`,
`subinterp_exec/2`, `subinterp_cast/4`, `subinterp_async_call/4`,
`subinterp_await/1,2`, and `subinterp_pool_*` are all gone. Use
`py_context:new(#{mode => owngil})` instead — it gives the same
parallelism with OTP supervision and automatic cleanup.
`py:subinterp_supported/0` (capability probe) and `py:parallel/1`
(which routes through the context API) stay.
- Internal `py_execution_mode_t` collapsed from 3 values to 2 (`free_threaded`
/ `gil`); `py_nif:execution_mode/0` returns `free_threaded | gil` instead
of the old `free_threaded | subinterp | multi_executor`.
- `examples/reactor_owngil_example.erl` deleted (called nonexistent
`py:subinterp_reactor_*` functions; pre-existing breakage).
## 2.3.1 (2026-04-01)
### Fixed
- **Executor affinity for numpy/torch** - Workers are now assigned a fixed executor
thread at creation. All calls from the same worker go to the same executor,
preventing thread state corruption in libraries like numpy and PyTorch that
have thread-local state. Fixes segfaults when using sentence-transformers
or other ML libraries.
## 2.3.0 (2026-03-29)
### Removed
- **ASGI/WSGI Support** - The `py_asgi` and `py_wsgi` modules have been removed
- `py_asgi:run/4,5` - ASGI application runner
- `py_wsgi:run/3,4` - WSGI application runner
- For web framework integration, use `py:call` with event loop contexts or the Channel API
- See [Migration Guide](docs/migration.md#asgiwsgi-modules-removed) for alternatives
### Added
- **SharedDict** - Process-scoped shared dictionaries for cross-process state
- `py:shared_dict_new/0` - Create a new SharedDict
- `py:shared_dict_get/2,3` - Get value with optional default
- `py:shared_dict_set/3` - Set key-value pair
- `py:shared_dict_del/2` - Delete a key
- `py:shared_dict_keys/1` - List all keys
- `py:shared_dict_destroy/1` - Explicit cleanup
- Python access via `erlang.SharedDict` with dict-like interface
- Mutex-protected for concurrent access (~300k ops/sec)
- Pickle serialization for complex types
- See [SharedDict documentation](docs/shared-dict.md) for details
## 2.2.0 (2026-03-24)
### Added
- **OWN_GIL Mode** - True parallel Python execution with Python 3.14+ subinterpreters
- Each subinterpreter runs with its own GIL (`Py_GIL_OWN`) in a dedicated thread
- Full isolation between interpreters (separate namespaces, modules, state)
- `py_context:start_link(N, owngil)` to create OWN_GIL contexts
- Enables true parallelism for CPU-bound Python workloads
- See [OWN_GIL Internals](docs/owngil_internals.md) for architecture details
- **Process-Bound Python Environments** - Per-Erlang-process Python namespaces
- Each Erlang process gets isolated Python globals/locals
- State persists across calls within the same process
- Automatic cleanup when Erlang process terminates
- See [Process-Bound Environments](docs/process-bound-envs.md) for details
- **Event Loop Pool** - Process affinity for parallel async execution
- `py_event_loop_pool` distributes async tasks across multiple event loops
- Scheduler-affinity routing for cache-friendly execution
- Supports worker, subinterp, and owngil modes
- **ByteChannel API** - Raw byte streaming without term serialization
- `py_byte_channel:new/0,1` - Create byte channels
- `py_byte_channel:send/2` - Send raw bytes
- `py_byte_channel:recv/1,2` - Receive bytes
- Python `ByteChannel` class with sync/async iteration
- Ideal for HTTP bodies, file streaming, binary protocols
- **PyBuffer API** - Zero-copy buffer for streaming input
- `py_buffer:new/0,1` - Create buffers with optional max size
- `py_buffer:write/2` - Write data to buffer
- Python `PyBuffer` class with file-like interface (`read`, `readline`, `readlines`)
- Non-blocking reads for async I/O patterns
- See [Buffer API](docs/buffer.md) for details
- **True streaming API** - New `py:stream_start/3,4` and `py:stream_cancel/1` functions
for event-driven streaming from Python generators. Unlike `py:stream/3,4` which
collects all values at once, `stream_start` sends `{py_stream, Ref, {data, Value}}`
messages as values are yielded. Supports both sync and async generators. Useful for
LLM token streaming, real-time data feeds, and processing large sequences incrementally.
- **`erlang.whereis(name)`** - Lookup registered Erlang process PIDs from Python
- Returns `erlang.Pid` object or `None` if not registered
- Enables Python code to discover and message named processes
- **`erlang.schedule_inline(callback)`** - Inline continuation scheduling
- Release dirty scheduler and continue with callback in same context
- Preserves globals/locals across the continuation
- Useful for cooperative long-running tasks
- **`py:spawn_call/3,4,5`** - Fire-and-forget with result delivery
- Executes Python call asynchronously
- Sends `{py_result, Ref, Result}` to caller when complete
- Non-blocking alternative to `py:call` for async patterns
- **Explicit bytes conversion** - `{bytes, Binary}` tuple for round-trip safety
- Erlang binaries convert to Python `str` by default
- Use `{bytes, Binary}` to force Python `bytes` type
- Ensures correct handling for binary protocols
- **Import caching API** - Lazy module import with caching
- `py:import/1,2` - Import and cache modules
- `py:add_import/1,2` - Register imports applied to all contexts
- `py:add_path/1` - Add to sys.path across all contexts
- Per-interpreter caching with generation tracking
- **Per-interpreter preload code** - Execute code in new interpreters
- Configure via `{erlang_python, [{preload_code, <<"import mylib">>}]}`
- Code runs with inherited globals from main interpreter
- Useful for initializing common imports/state
### Fixed
- **Channel notification for create_task** - Fixed async channel receive hanging when using
`py_event_loop:create_task`. The `event_loop_add_pending()` now sends `task_ready` to the
worker, not just `pthread_cond_signal`. Also fixed Python 3.9 compatibility in ByteChannel
(`Optional[bytes]` instead of `bytes | None`)
- **Channel waiter race condition** - Fixed `waiter_exists` errors during fast async iteration.
Waiter state is now cleared before releasing mutex, preventing race where callback fires
before `channel_send` clears `has_waiter`
- **Event Loop Isolation and Resource Safety** - Three fixes for event loop and atom handling
- **Single-loop-per-interpreter enforcement** - Prevents multiple `ErlangEventLoop` instances
from causing event confusion. Added `_has_loop_ref()` check that detects running loops;
attempting to create a second loop while one is running raises `RuntimeError`
- **Atom creation safety** - Added Python-level caching with configurable limit (10000 default,
`ERLANG_PYTHON_MAX_ATOMS` env var) to prevent BEAM atom table exhaustion from untrusted code.
The `erlang.atom()` API now goes through the cached wrapper; internal `_atom()` NIF still available
- **Global capsule resource leak** - Added `global_loop_capsule_destructor` that properly calls
`enif_release_resource()` when capsule is garbage collected. Previously NULL destructor caused
reference leaks on each `ErlangEventLoop` creation
- **Python 3.14 venv activation** - Fixed `.pth` file processing in subinterpreters. Python 3.14
stricter module isolation prevented `sys._venv_site_packages` from persisting across eval/exec calls.
Now embeds site-packages path directly in the exec code string
- **OWN_GIL Safety Fixes** - Critical fixes for OWN_GIL subinterpreter mode
- **Mutex leak in erlang module** - `async_futures_mutex` now always destroyed in
`erlang_module_free()` regardless of `pipe_initialized` flag
- **ABBA deadlock prevention** - Fixed lock ordering in `event_loop_down()` and
`event_loop_destructor()` to acquire GIL before `namespaces_mutex`, matching the
normal execution path and preventing deadlocks
- **Dangling env pointer detection** - Added `interp_id` validation in
`owngil_execute_*_with_env()` functions to detect and reject env resources
created by a different interpreter, returning `{error, env_wrong_interpreter}`
- **OWN_GIL callback documentation** - Documented that `erlang.call()` from OWN_GIL
contexts uses `thread_worker_call()` rather than suspension/resume protocol;
re-entrant calls to the same OWN_GIL context are not supported
### Changed
- **`py:cast` is now fire-and-forget** - `py:cast/3,4,5` no longer returns a reference.
For async calls with result delivery, use the new `py:spawn_call/3,4,5` instead.
- **OWN_GIL requires Python 3.14+** - The OWN_GIL subinterpreter mode requires Python 3.14
or later due to C extension compatibility issues in earlier versions. Use `worker` or
`subinterp` modes for Python 3.12-3.13.
- **Removed auto-started io pool** - The io pool is no longer started automatically at
application startup to reduce memory usage. Users who need a dedicated I/O pool can
create one manually via `py_context_router:start_pool(io, 10, worker)`. The configuration
options `io_pool_size` and `io_pool_mode` have been removed.
- **Removed py_event_router** - Removed legacy `py_event_router` module. The `py_event_worker`
now handles all event loop functionality including FD events, timers, and task processing.
This simplifies the architecture by consolidating event handling into a single worker process.
The `py_nif:set_shared_router/1` function has been removed.
- **Config-based initialization** - Import and path configuration via application environment
- Configure imports: `{erlang_python, [{imports, [{json, dumps}]}]}`
- Configure paths: `{erlang_python, [{paths, ["/path/to/modules"]}]}`
- Applied immediately to all running interpreters
- See [Imports documentation](docs/imports.md) for details
### Performance
- **Direct NIF channel operations** - Channel send/receive bypass `erlang.call()` overhead
for up to 1760x speedup in raw throughput benchmarks
- **nif_process_ready_tasks optimization** - ~15% improvement in async task processing
- Replace `asyncio.iscoroutine()` with `PyCoro_CheckExact` C API
- Use stack buffers for module/func strings
- Cache `asyncio.events` module
- Pool `ErlNifEnv` allocations with mutex protection
## 2.1.0 (2026-03-12)
### Added
- **Async Task API** - uvloop-inspired task submission from Erlang
- `py_event_loop:run/3,4` - Blocking run of async Python functions
- `py_event_loop:create_task/3,4` - Non-blocking task submission with reference
- `py_event_loop:await/1,2` - Wait for task result with timeout
- `py_event_loop:spawn_task/3,4` - Fire-and-forget task execution
- Thread-safe submission via `enif_send` (works from dirty schedulers)
- Message-based result delivery via `{async_result, Ref, Result}`
- See [Async Task API docs](docs/asyncio.md#async-task-api-erlang) for details
- **`erlang.spawn_task(coro)`** - Spawn async tasks from both sync and async contexts
- Works in sync code called by Erlang (where `asyncio.get_running_loop()` fails)
- Returns `asyncio.Task` for optional await/cancel (fire-and-forget pattern)
- Automatically wakes up the event loop in sync context
- **Explicit Scheduling API** - Control dirty scheduler release from Python
- `erlang.schedule(callback, *args)` - Release scheduler, continue via Erlang callback
- `erlang.schedule_py(module, func, args, kwargs)` - Release scheduler, continue in Python
- `erlang.consume_time_slice(percent)` - Check if NIF time slice exhausted
- `ScheduleMarker` type for cooperative long-running tasks
- See [Scheduling API docs](docs/asyncio.md#explicit-scheduling-api)
- **Distributed Python Execution** - Documentation and Docker demo
- Run Python across Erlang nodes using `rpc:call`
- Docker Compose setup for testing distributed patterns
- See [Distributed Execution docs](docs/distributed.md)
### Changed
- **Event Loop Performance Optimizations**
- Growable pending queue with capacity doubling (256 to 16384)
- Snapshot-detach pattern to reduce mutex contention
- Callable cache (64 slots) avoids PyImport/GetAttr per task
- Task wakeup coalescing with atomic flag
- Drain-until-empty loop for faster task processing
### Fixed
- `ensure_venv` now always installs dependencies, even if venv exists
- `erlang.sleep()` timing in sync context
- `time()` returns fresh value when loop not running
- Handle pooling bugs in ErlangEventLoop
- Task wakeup race causing batch task stalls
## 2.0.0 (2026-03-09)
### Added
- **Virtual Environment Management** - Automatic venv creation and activation
- `py:ensure_venv/2,3` - Create venv if missing, then activate
- Automatically detects Python executable
- Supports pip install of dependencies
- **File Descriptor Duplication** - Safe socket handoff from Erlang to Python
- `py:dup_fd/1` - Duplicate fd for independent ownership
- Prevents double-close issues when passing sockets to Python reactor
- **Custom Pool Support** - Create pools on demand for CPU-bound and I/O-bound operations
- `default` pool - Automatically started, sized to number of schedulers
- `py_context_router:start_pool/2,3` - Start named pools programmatically
- `py_context_router:stop_pool/1` - Stop a named pool
- `py_context_router:pool_started/1` - Check if a pool is running
- `py_context_router:get_context(Pool)` - Get context from a named pool
- `py_context_router:num_contexts(Pool)` - Get pool size
- `py_context_router:contexts(Pool)` - Get all contexts in a pool
- `py_context_router:lookup_pool(Module, Func)` - Query pool routing
- `py:call(PoolName, Module, Func, Args)` - Execute on a specific pool
- Registration-based routing (no call site changes needed):
- `py:register_pool(io, requests)` - Route all `requests.*` calls to io pool
- `py:register_pool(io, {aiohttp, get})` - Route specific function to io pool
- `py:unregister_pool(Module)` - Remove module registration
- `py:unregister_pool({Module, Func})` - Remove function registration
- Automatic routing: `py:call(requests, get, [Url])` goes to io pool when registered
- Backward compatible: existing code using `py:call/3,4,5` works unchanged
- New test suite: `test/py_pool_SUITE.erl`
- **Channel API** - Bidirectional message passing between Erlang and Python
- `py_channel:new/0,1` - Create channels with optional backpressure (`max_size`)
- `py_channel:send/2` - Send Erlang terms to Python (returns `busy` on backpressure)
- `py_channel:close/1` - Close channel, signals `StopIteration` to Python
- Python `Channel` class with sync and async interfaces:
- `channel.receive()` - Blocking receive (suspends Python, yields to Erlang)
- `channel.try_receive()` - Non-blocking receive
- `await channel.async_receive()` - Asyncio-compatible receive
- `for msg in channel:` - Sync iteration
- `async for msg in channel:` - Async iteration
- `erlang.channel.reply(pid, term)` - Send messages to Erlang processes
- Zero-copy IOQueue buffering via `enif_ioq`
- 8x faster than Reactor for small messages, 2x faster for 16KB messages
- **OWN_GIL Subinterpreter Thread Pool** - True parallelism with Python 3.12+ subinterpreters
- Each subinterpreter runs in its own thread with its own GIL (`Py_GIL_OWN`)
- Thread pool manages N subinterpreters for parallel Python execution
- `py:context(N)` returns the Nth context PID for explicit context selection
- `py_context_router` provides scheduler-affinity routing for automatic distribution
- Cast operations are 25-30% faster compared to worker mode
- Full isolation between subinterpreters (separate namespaces, modules, state)
- New C files: `py_subinterp_pool.c`, `py_subinterp_pool.h`
- **`erlang.reactor` module** - FD-based protocol handling for building custom servers
- `reactor.Protocol` - Base class for implementing protocols
- `reactor.serve(sock, protocol_factory)` - Serve connections using a protocol
- `reactor.run_fd(fd, protocol_factory)` - Handle a single FD with a protocol
- Integrates with Erlang's `enif_select` for efficient I/O multiplexing
- Zero-copy buffer management for high-throughput scenarios
- Supports SHARED_GIL subinterpreters via `py_reactor_context`
- Each reactor context has isolated protocol factory when using `mode=subinterp`
- **ETF encoding for PIDs and References** - Full Erlang term format support
- Erlang PIDs encode/decode properly in ETF binary format
- Erlang References encode/decode properly in ETF binary format
- Enables proper serialization for distributed Erlang communication
- **PID serialization** - Erlang PIDs now convert to `erlang.Pid` objects in Python
and back to real PIDs when returned to Erlang. Previously, PIDs fell through to
`None` (Erlang→Python) or string representation (Python→Erlang).
- **`erlang.send(pid, term)`** - Fire-and-forget message passing from Python to
Erlang processes. Uses `enif_send()` directly with no suspension or blocking.
Raises `erlang.ProcessError` if the target process is dead.
- **`erlang.ProcessError`** - New exception for dead/unreachable process errors.
Subclass of `Exception`, so it's catchable with `except Exception` or
`except erlang.ProcessError`.
- **Audit hook sandbox** - Block dangerous operations when running inside Erlang VM
- Uses Python's `sys.addaudithook()` (PEP 578) for low-level blocking
- Blocks: `os.fork`, `os.system`, `os.popen`, `os.exec*`, `os.spawn*`, `subprocess.Popen`
- Raises `RuntimeError` with clear message about using Erlang ports instead
- Automatically installed when `py_event_loop` NIF is available
- **Process-per-context architecture** - Each Python context runs in dedicated process
- `py_context_process` - Gen_server managing a single Python context
- `py_context_sup` - Supervisor for context processes
- `py_context_router` - Routes calls to appropriate context process
- Improved isolation between contexts
- Better crash recovery and resource management
- **Worker thread pool** - High-throughput Python operations
- Configurable pool size for parallel execution
- Efficient work distribution across threads
- **`py:contexts_started/0`** - Helper to check if contexts are ready
### Changed
- **`py:call_async` renamed to `py:cast`** - Follows gen_server convention where
`call` is synchronous and `cast` is asynchronous. The semantics are identical,
only the name changed.
- **Unified `erlang` Python module** - Consolidated callback and event loop APIs
- `erlang.run(coro)` - Run coroutine with ErlangEventLoop (like uvloop.run)
- `erlang.new_event_loop()` - Create new ErlangEventLoop instance
- `erlang.install()` - Install ErlangEventLoopPolicy (deprecated in 3.12+)
- `erlang.EventLoopPolicy` - Alias for ErlangEventLoopPolicy
- Removed separate `erlang_asyncio` module - all functionality now in `erlang`
- **Async worker backend replaced with event loop model** - The pthread+usleep
polling async workers have been replaced with an event-driven model using
`py_event_loop` and `enif_select`:
- Removed `py_async_worker.erl` and `py_async_worker_sup.erl`
- Removed `py_async_worker_t` and `async_pending_t` structs from C code
- Deprecated `async_worker_new`, `async_call`, `async_gather`, `async_stream` NIFs
- Added `py_event_loop_pool.erl` for managing event loop-based async execution
- Added `py_event_loop:run_async/2` for submitting coroutines to event loops
- Added `nif_event_loop_run_async` NIF for direct coroutine submission
- Added `_run_and_send` wrapper in Python for result delivery via `erlang.send()`
- **Internal change**: `py:async_call/3,4` and `py:await/1,2` API unchanged
- **`SuspensionRequired` base class** - Now inherits from `BaseException` instead
of `Exception`. This prevents ASGI/WSGI middleware `except Exception` handlers
from intercepting the suspension control flow used by `erlang.call()`.
- **Per-interpreter isolation in py_event_loop.c** - Removed global state for
proper subinterpreter support. Each interpreter now has isolated event loop state.
- **ErlangEventLoopPolicy always returns ErlangEventLoop** - Previously only
returned ErlangEventLoop for main thread; now consistent across all threads.
### Deprecated
- **`py_asgi` module** - Deprecated in favor of the Channel API (`py_channel`)
or Reactor API (`erlang.reactor`). The module still works but will be removed
in a future release.
- **`py_wsgi` module** - Deprecated in favor of the Channel API (`py_channel`)
or Reactor API (`erlang.reactor`). The module still works but will be removed
in a future release.
### Removed
- **Context affinity functions** - Removed `py:bind`, `py:unbind`, `py:is_bound`,
`py:with_context`, and `py:ctx_*` functions. The new `py_context_router` provides
automatic scheduler-affinity routing. For explicit context control, use
`py_context_router:bind_context/1` and `py_context:call/5`.
- **Signal handling support** - Removed `add_signal_handler`/`remove_signal_handler`
from ErlangEventLoop. Signal handling should be done at the Erlang VM level.
Methods now raise `NotImplementedError` with guidance.
- **Subprocess support** - ErlangEventLoop raises `NotImplementedError` for
`subprocess_shell` and `subprocess_exec`. Use Erlang ports (`open_port/2`)
for subprocess management instead.
### Fixed
- **`py_reactor_context` now extends erlang module in subinterpreters** - Previously,
`py_reactor_context` with `mode=subinterp` would fail to import `erlang.reactor`
because the erlang module extension was not applied. Now calls
`py_context:extend_erlang_module_in_context/1` after context creation.
- **FD stealing and UDP connected socket issues** - Fixed file descriptor handling
for UDP sockets in connected mode
- **Context test expectations** - Updated tests for Python contextvars behavior
- **Unawaited coroutine warnings** - Fixed warnings in test suite
- **Timer scheduling for standalone ErlangEventLoop** - Fixed timer callbacks not
firing for loops created outside the main event loop infrastructure
- **Subinterpreter cleanup and thread worker re-registration** - Fixed cleanup
issues when subinterpreters are destroyed and recreated
- **ProcessError exception class identity in subinterpreters** - Fixed exception
class mismatch when raising `erlang.ProcessError` in subinterpreter contexts.
The exception class is now looked up from the current interpreter's `erlang`
module at runtime instead of using a global variable.
- **Thread worker handlers not re-registering after app restart** - Workers now
properly re-register when application restarts
- **Timeout handling** - Improved timeout handling across the codebase
- **Eval locals_term initialization** - Fixed uninitialized variable in eval
- **Two race conditions in worker pool** - Fixed concurrent access issues
- **`activate_venv/1` now processes `.pth` files** - Uses `site.addsitedir()` instead of
`sys.path.insert()` so that editable installs (uv, pip -e, poetry) work correctly.
New paths are moved to the front of `sys.path` for proper priority.
- **`deactivate_venv/0` now restores `sys.path`** - The previous implementation used
`py:eval` with semicolon-separated statements which silently failed (eval only accepts
expressions). Switched to `py:exec` for correct statement execution.
### Performance
- **Async coroutine latency reduced from ~10-20ms to <1ms** - The event loop model
eliminates pthread polling overhead
- **Zero CPU usage when idle** - Event-driven instead of usleep-based polling
- **No extra threads** - Coroutines run on the existing event loop infrastructure
## 1.8.1 (2026-02-25)
### Fixed
- **ASGI scope caching bug** - HTTP method was not treated as a dynamic field in the
scope template cache. This caused incorrect method values when the same path was
accessed with different HTTP methods (e.g., GET /path followed by POST /path would
return method="GET" for both requests).
## 1.8.0 (2026-02-25)
### Added
- **ASGI NIF Optimizations** - Six optimizations for high-performance ASGI request handling
- **Direct Response Tuple Extraction** - Extract `(status, headers, body)` directly without generic conversion
- **Pre-Interned Header Names** - 16 common HTTP headers cached as PyBytes objects
- **Cached Status Code Integers** - 14 common HTTP status codes cached as PyLong objects
- **Zero-Copy Request Body** - Large bodies (≥1KB) use buffer protocol for zero-copy access
- **Scope Template Caching** - Thread-local cache of 64 scope templates keyed by path hash
- **Lazy Header Conversion** - Headers converted on-demand for requests with ≥4 headers
- **erlang_asyncio Module** - Asyncio-compatible primitives using Erlang's native scheduler
- `erlang_asyncio.sleep(delay, result=None)` - Sleep using Erlang's `erlang:send_after/3`
- `erlang_asyncio.run(coro)` - Run coroutine with ErlangEventLoop
- `erlang_asyncio.gather(*coros)` - Run coroutines concurrently
- `erlang_asyncio.wait_for(coro, timeout)` - Wait with timeout
- `erlang_asyncio.wait(fs, timeout, return_when)` - Wait for multiple futures
- `erlang_asyncio.create_task(coro)` - Create background task
- `erlang_asyncio.ensure_future(coro)` - Wrap coroutine in Future
- `erlang_asyncio.shield(arg)` - Protect from cancellation
- `erlang_asyncio.timeout` - Context manager for timeouts
- Event loop functions: `get_event_loop()`, `new_event_loop()`, `set_event_loop()`, `get_running_loop()`
- Re-exports: `TimeoutError`, `CancelledError`, `ALL_COMPLETED`, `FIRST_COMPLETED`, `FIRST_EXCEPTION`
- **Erlang Sleep NIF** - Synchronous sleep primitive for Python
- `py_event_loop._erlang_sleep(delay_ms)` - Sleep using Erlang timer
- Releases GIL during sleep, no Python event loop overhead
- Uses pthread condition variables for efficient blocking
- `py_nif:dispatch_sleep_complete/2` - NIF to signal sleep completion
- **Scalable I/O Model** - Worker-per-context architecture
- `py_event_worker` - Dedicated worker process per Python context
- Combined FD event dispatch and reselect via `handle_fd_event_and_reselect` NIF
- Sleep tracking with `sleeps` map in worker state
- **New Test Suite** - `test/py_erlang_sleep_SUITE.erl` with 8 tests
- `test_erlang_sleep_available` - Verify NIF is exposed
- `test_erlang_sleep_basic` - Basic functionality
- `test_erlang_sleep_zero` - Zero delay returns immediately
- `test_erlang_sleep_accuracy` - Timing accuracy
- `test_erlang_asyncio_module` - Module functions present
- `test_erlang_asyncio_gather` - Concurrent execution
- `test_erlang_asyncio_wait_for` - Timeout support
- `test_erlang_asyncio_create_task` - Background tasks
### Performance
- **ASGI marshalling optimizations** - 40-60% improvement for typical ASGI workloads
- Direct response extraction: 5-10% improvement
- Pre-interned headers: 3-5% improvement
- Cached status codes: 1-2% improvement
- Zero-copy body buffers: 10-15% for large bodies (≥1KB)
- Scope template caching: 15-20% for repeated paths
- Lazy header conversion: 5-10% for apps accessing few headers
- **Eliminates event loop overhead** for sleep operations (~0.5-1ms saved per call)
- **Sub-millisecond timer precision** via BEAM scheduler (vs 10ms asyncio polling)
- **Zero CPU when idle** - event-driven, no polling
## 1.7.1 (2026-02-23)
### Fixed
- **Hex package missing priv directory** - Added explicit `files` configuration to include
`priv/erlang_loop.py` and other necessary files in the hex.pm package
## 1.7.0 (2026-02-23)
### Added
- **Shared Router Architecture for Event Loops**
- Single `py_event_router` process handles all event loops
- Timer and FD messages include loop identity for correct dispatch
- Eliminates need for per-loop router processes
- Handle-based Python C API using PyCapsule for loop references
- **Per-Loop Capsule Architecture** - Each `ErlangEventLoop` instance has its own isolated capsule
- Dedicated pending queue per loop for proper event routing
- Full asyncio support (timers, FD operations) with correct loop isolation
- Safe for multi-threaded Python applications where each thread needs its own loop
- See `docs/asyncio.md` for usage and architecture details
## 1.6.1 (2026-02-22)
### Fixed
- **ASGI headers now correctly use bytes instead of str** - Fixed ASGI spec compliance
issue where headers were being converted to Python `str` objects instead of `bytes`.
The ASGI specification requires headers to be `list[tuple[bytes, bytes]]`. This was
causing authentication failures and form parsing issues with frameworks like Starlette
and FastAPI, which search for headers using bytes keys (e.g., `b"content-type"`).
- Added explicit header handling in `asgi_scope_from_map()` to bypass generic conversion
- Headers are now correctly converted using `PyBytes_FromStringAndSize()`
- Supports both list `[name, value]` and tuple `{name, value}` header formats from Erlang
- Fixes GitHub issue #1
## 1.6.0 (2026-02-22)
### Added
- **Python Logging Integration** - Forward Python's `logging` module to Erlang's `logger`
- `py:configure_logging/0,1` - Setup Python logging to forward to Erlang
- `erlang.ErlangHandler` - Python logging handler that sends to Erlang
- `erlang.setup_logging(level, format)` - Configure logging from Python
- Fire-and-forget architecture using `enif_send()` for non-blocking messaging
- Level filtering at NIF level for performance (skip message creation for filtered logs)
- Log metadata includes module, line number, and function name
- Thread-safe - works from any Python thread
- **Distributed Tracing** - Collect trace spans from Python code
- `py:enable_tracing/0`, `py:disable_tracing/0` - Enable/disable span collection
- `py:get_traces/0` - Retrieve collected spans
- `py:clear_traces/0` - Clear collected spans
- `erlang.Span(name, **attrs)` - Context manager for creating spans
- `erlang.trace(name)` - Decorator for tracing functions
- Span events via `span.event(name, **attrs)`
- Automatic parent/child span linking via thread-local storage
- Error status capture with exception details
- Duration tracking in microseconds
- **New Erlang modules**
- `py_logger` - gen_server receiving log messages from Python workers
- `py_tracer` - gen_server collecting and managing trace spans
- **New C source**
- `c_src/py_logging.c` - NIF implementations for logging and tracing
- **Documentation and examples**
- `docs/logging.md` - Logging and tracing documentation
- `examples/logging_example.erl` - Working escript example
- Updated `docs/getting-started.md` with logging/tracing section
- **New test suite**
- `test/py_logging_SUITE.erl` - 9 tests for logging and tracing
- `ATOM_NIL` for Elixir `nil` compatibility in type conversions
### Performance
- **Type conversion optimizations** - Faster Python ↔ Erlang marshalling
- Use `enif_is_identical` for atom comparison instead of `strcmp`
- Use `PyLong_AsLongLongAndOverflow` to avoid exception machinery
- Cache `numpy.ndarray` type at init for fast isinstance checks
- Stack allocate small tuples/maps (≤16 elements) to avoid heap allocation
- Use `enif_make_map_from_arrays` for O(n) map building vs O(n²) puts
- Reorder type checks for web workloads (strings/dicts first)
- UTF-8 decode with bytes fallback for invalid sequences
- **Fire-and-forget NIF architecture** - Log and trace calls never block Python execution
- Uses `enif_send()` to dispatch messages asynchronously to Erlang processes
- Python code continues immediately after sending, no round-trip wait
- **NIF-level log filtering** - Messages below threshold are discarded before term creation
- Volatile bool flags for O(1) receiver availability checks
- Level threshold stored in C global, no Erlang callback needed
- **Minimal term allocation** - Direct Erlang term building without intermediate structures
- Timestamps captured at NIF level using `enif_monotonic_time()`
### Fixed
- **Python 3.12+ event loop thread isolation** - Fixed asyncio timeouts on Python 3.12+
- `ErlangEventLoop` now only used for main thread; worker threads get `SelectorEventLoop`
- Async worker threads bypass the policy to create `SelectorEventLoop` directly
- Per-call `ErlNifEnv` for thread-safe timer scheduling in free-threaded mode
- Fail-fast error handling in `erlang_loop.py` instead of silent hangs
- Added `gil_acquire()`/`gil_release()` helpers to avoid GIL double-acquisition
## 1.5.0 (2026-02-18)
### Added
- **`py_asgi` module** - Optimized ASGI request handling with:
- Pre-interned Python string keys (15+ ASGI scope keys)
- Cached constant values (http type, HTTP versions, methods, schemes)
- Thread-local response pooling (16 slots per thread, 4KB initial buffer)
- Direct NIF path bypassing generic py:call()
- ~60-80% throughput improvement over py:call()
- Configurable runner module via `runner` option
- Sub-interpreter and free-threading (Python 3.13+) support
- **`py_wsgi` module** - Optimized WSGI request handling with:
- Pre-interned WSGI environ keys
- Direct NIF path for marshalling
- ~60-80% throughput improvement over py:call()
- Sub-interpreter and free-threading support
- **Web frameworks documentation** - New documentation at `docs/web-frameworks.md`
## 1.4.0 (2026-02-18)
### Added
- **Erlang-native asyncio event loop** - Custom asyncio event loop backed by Erlang's scheduler
- `ErlangEventLoop` class in `priv/erlang_loop.py`
- Sub-millisecond latency via Erlang's `enif_select` (vs 10ms polling)
- Zero CPU usage when idle - no busy-waiting or polling overhead
- Full GIL release during waits for better concurrency
- Native Erlang scheduler integration for I/O events
- Event loop policy via `get_event_loop_policy()`
- **TCP support for asyncio event loop**
- `create_connection()` - TCP client connections
- `create_server()` - TCP server with accept loop
- `_ErlangSocketTransport` - Non-blocking socket transport with write buffering
- `_ErlangServer` - TCP server with `serve_forever()` support
- **UDP/datagram support for asyncio event loop**
- `create_datagram_endpoint()` - Create UDP endpoints with full parameter support
- `_ErlangDatagramTransport` - Datagram transport implementation
- Parameters: `local_addr`, `remote_addr`, `reuse_address`, `reuse_port`, `allow_broadcast`
- `DatagramProtocol` callbacks: `datagram_received()`, `error_received()`
- Support for both connected and unconnected UDP
- New NIF helpers: `create_test_udp_socket`, `sendto_test_udp`, `recvfrom_test_udp`, `set_udp_broadcast`
- New test suite: `test/py_udp_e2e_SUITE.erl`
- **Asyncio event loop documentation**
- New documentation: `docs/asyncio.md`
- Updated `docs/getting-started.md` with link to asyncio documentation
### Performance
- **Event loop optimizations**
- Fixed `run_until_complete` callback removal bug (was using two different lambda references)
- Cached `ast.literal_eval` lookup at module initialization (avoids import per callback)
- O(1) timer cancellation via handle-to-callback_id reverse map (was O(n) iteration)
- Detach pending queue under mutex, build Erlang terms outside lock (reduced contention)
- O(1) duplicate event detection using hash set (was O(n) linear scan)
- Added `PERF_BUILD` cmake option for aggressive optimizations (-O3, LTO, -march=native)
## 1.3.2 (2026-02-17)
### Fixed
- **torch/PyTorch introspection compatibility** - Fixed `AttributeError: 'erlang.Function'
object has no attribute 'endswith'` when importing torch or sentence_transformers in
contexts where erlang_python callbacks are registered.
- Root cause: torch does dynamic introspection during import, iterating through Python's
namespace and calling `.endswith()` on objects. The `erlang` module's `__getattr__` was
returning `ErlangFunction` wrappers for *any* attribute access.
- Solution: Added C-side callback name registry. Now `__getattr__` only returns
`ErlangFunction` wrappers for actually registered callbacks. Unregistered attributes
raise `AttributeError` (normal Python behavior).
- New test: `test_callback_name_registry` in `py_reentrant_SUITE.erl`
## 1.3.1 (2026-02-16)
### Fixed
- **Hex.pm packaging** - Added `files` section to app.src to include build scripts
(`do_cmake.sh`, `do_build.sh`) and other necessary files in the hex.pm package
## 1.3.0 (2026-02-16)
### Added
- **Asyncio Support** - New `erlang.async_call()` for asyncio-compatible callbacks
- `await erlang.async_call('func', arg1, arg2)` - Call Erlang from async Python code
- Integrates with asyncio event loop via `add_reader()`
- No exceptions raised for control flow (unlike `erlang.call()`)
- Releases dirty NIF thread while waiting (non-blocking)
- Works with FastAPI, Starlette, aiohttp, and other ASGI frameworks
- Supports concurrent calls via `asyncio.gather()`
- New test: `test_async_call` in `py_reentrant_SUITE.erl`
- New test module: `test/py_test_async.py`
- Updated documentation: `docs/threading.md` - Added Asyncio Support section
### Fixed
- **Flag-based callback detection in replay path** - Fixed SuspensionRequired exceptions
leaking when ASGI middleware catches and re-raises exceptions. The replay path in
`nif_resume_callback_dirty` now uses flag-based detection (checking `tl_pending_callback`)
instead of exception-type detection.
### Changed
- **C code optimizations and refactoring**
- **Thread safety fixes**: Used `pthread_once` for async callback initialization,
fixed mutex held during Python calls in async event loop thread
- **Timeout handling**: Added `read_with_timeout()` and `read_length_prefixed_data()`
helpers with proper timeouts on all blocking pipe reads (30s for callbacks, 10s for spawns)
- **Code deduplication**: Merged `create_suspended_state()` and
`create_suspended_state_from_existing()` into unified `create_suspended_state_ex()`,
extracted `build_pending_callback_exc_args()` and `build_suspended_result()` helpers
- **Performance**: Optimized list conversion using `enif_make_list_cell()` to build
lists directly without temporary array allocation
- Removed unused `make_suspended_term()` function
## 1.2.0 (2026-02-15)
### Added
- **Context Affinity** - Bind Erlang processes to dedicated Python workers for state persistence
- `py:bind()` / `py:unbind()` - Bind current process to a worker, preserving Python state
- `py:bind(new)` - Create explicit context handles for multiple contexts per process
- `py:with_context(Fun)` - Scoped helper with automatic bind/unbind
- Context-aware functions: `py:ctx_call/4-6`, `py:ctx_eval/2-4`, `py:ctx_exec/2`
- Automatic cleanup via process monitors when bound processes die
- O(1) ETS-based binding lookup for minimal overhead
- New test suite: `test/py_context_SUITE.erl`
- **Python Thread Support** - Any spawned Python thread can now call `erlang.call()` without blocking
- Supports `threading.Thread`, `concurrent.futures.ThreadPoolExecutor`, and any other Python threads
- Each spawned thread lazily acquires a dedicated "thread worker" channel
- One lightweight Erlang process per Python thread handles callbacks
- Automatic cleanup when Python thread exits via `pthread_key_t` destructor
- New module: `py_thread_handler.erl` - Coordinator and per-thread handlers
- New C file: `py_thread_worker.c` - Thread worker pool management
- New test suite: `test/py_thread_callback_SUITE.erl`
- New documentation: `docs/threading.md` - Threading support guide
- **Reentrant Callbacks** - Python→Erlang→Python callback chains without deadlocks
- Exception-based suspension mechanism interrupts Python execution cleanly
- Callbacks execute in separate processes to prevent worker pool exhaustion
- Supports arbitrarily deep nesting (tested up to 10+ levels)
- Transparent to users - `erlang.call()` works the same, just without deadlocks
- New test suite: `test/py_reentrant_SUITE.erl`
- New examples: `examples/reentrant_demo.erl` and `examples/reentrant_demo.py`
### Changed
- Callback handlers now spawn separate processes for execution, allowing workers
to remain available for nested `py:eval`/`py:call` operations
- **Modular C code structure** - Split monolithic `py_nif.c` (4,335 lines) into
logical modules for better maintainability:
- `py_nif.h` - Shared header with types, macros, and declarations
- `py_convert.c` - Bidirectional type conversion (Python ↔ Erlang)
- `py_exec.c` - Python execution engine and GIL management
- `py_callback.c` - Erlang callback support and asyncio integration
- Uses `#include` approach for single compilation unit (no build changes needed)
### Fixed
- **Multiple sequential erlang.call()** - Fixed infinite loop when Python code makes
multiple sequential `erlang.call()` invocations in the same function. The replay
mechanism now falls back to blocking pipe behavior for subsequent calls after the
first suspension, preventing the infinite replay loop.
- **Memory safety in C NIF** - Fixed memory leaks and added NULL checks
- `nif_async_worker_new`: msg_env now freed on pipe/thread creation failure
- `multi_executor_stop`: shutdown requests now properly freed after join
- `create_suspended_state`: binary allocations cleaned up on failure paths
- Added NULL checks on all `enif_alloc_resource` and `enif_alloc_env` calls
- **Dialyzer warnings** - Added `{suspended, ...}` return type to NIF specs for
`worker_call`, `worker_eval`, and `resume_callback` functions
- **Dead code removal** - Cleaned up unused code discovered during code review:
- Removed `execute_direct()` function in `py_exec.c` (duplicated inline logic)
- Removed unused `ref` field from `async_pending_t` struct in `py_nif.h`
- Removed `worker_recv/2` from `py_nif.erl` (declared but never implemented in C)
### Documentation
- **Doxygen-style C documentation** - Added documentation to all C source files:
- Architecture overview with execution mode diagrams
- Type mapping tables for conversions
- GIL management patterns and best practices
- Suspension/resume flow diagrams for callbacks
- Function-level `@param`, `@return`, `@pre`, `@warning`, `@see` annotations
## 1.1.0 (2026-02-15)
### Added
- **Shared State API** - ETS-backed storage for sharing data between Python workers
- `state_set/get/delete/keys/clear` accessible from Python via `from erlang import ...`
- `py:state_store/fetch/remove/keys/clear` from Erlang
- Atomic counters with `state_incr/decr` (Python) and `py:state_incr/decr` (Erlang)
- New example: `examples/shared_state_example.erl`
- **Native Python Import Syntax** for Erlang callbacks
- `from erlang import my_func; my_func(args)` - most Pythonic
- `erlang.my_func(args)` - attribute-style access
- `erlang.call('my_func', args)` - legacy syntax still works
- **Module Reload** - Reload Python modules across all workers during development
- `py:reload(module)` uses `importlib.reload()` to refresh modules from disk
- `py_pool:broadcast` for sending requests to all workers
- **Documentation improvements**
- Added shared state section to getting-started, scalability, and ai-integration guides
- Added embedding caching example using shared state
- Added hex.pm badges to README
### Fixed
- **Memory safety** - Added NULL checks to all `enif_alloc()` calls in NIF code
- **Worker resilience** - Fixed crash in `py_subinterp_pool:terminate` when workers undefined
- **Streaming example** - Fixed to work with worker pool design (workers don't share namespace)
- **ETS table ownership** - Moved `py_callbacks` table creation to supervisor for resilience
### Changed
- Created `py_util` module to consolidate duplicate code (`to_binary/1`, `send_response/3`, `normalize_timeout/1-2`)
- Consolidated `async_await/2` to call `await/2` reducing duplication
## 1.0.0 (2026-02-14)
Initial release of erlang_python - Execute Python from Erlang/Elixir using dirty NIFs.
### Features
- **Python Integration**
- Call Python functions with `py:call/3-5`
- Evaluate expressions with `py:eval/1-3`
- Execute statements with `py:exec/1-2`
- Stream from Python generators with `py:stream/3-4`
- **Multiple Execution Modes** (auto-detected)
- Free-threaded Python 3.13+ (no GIL, true parallelism)
- Sub-interpreters Python 3.12+ (per-interpreter GIL)
- Multi-executor for older Python versions
- **Worker Pools**
- Main worker pool for synchronous calls
- Async worker pool for asyncio coroutines
- Sub-interpreter pool for parallel execution
- **Erlang/Elixir Callbacks**
- Register functions callable from Python via `py:register_function/2-3`
- Python code calls back with `erlang.call('name', args...)`
- **Virtual Environment Support**
- Activate venvs with `py:activate_venv/1`
- Use isolated package dependencies
- **Rate Limiting**
- ETS-based semaphore prevents overload
- Configurable max concurrent operations
- **Type Conversion**
- Automatic conversion between Erlang and Python types
- Integers, floats, strings, lists, tuples, maps/dicts, booleans
- **Memory Management**
- Access Python GC stats with `py:memory_stats/0`
- Force garbage collection with `py:gc/0-1`
- Memory tracing with `py:tracemalloc_start/stop`
### Examples
- `semantic_search.erl` - Text embeddings and similarity search
- `rag_example.erl` - Retrieval-Augmented Generation with Ollama
- `ai_chat.erl` - Interactive LLM chat
- `erlang_concurrency.erl` - 10x speedup with BEAM processes
- `elixir_example.exs` - Full Elixir integration demo
### Documentation
- Getting Started guide
- AI Integration guide
- Type Conversion reference
- Scalability and performance tuning
- Streaming with generators