Packages

Leak-free external process execution: process-group kill on timeout or BEAM death via a precompiled Rust shim

Current section

Files

Jump to
forcola lib forcola.ex
Raw

lib/forcola.ex

defmodule Forcola do
@moduledoc """
Leak-free external process execution.
Every command runs under a small Rust shim (a port program, not a NIF)
that places the child in its own process group via `setsid` and kills
the whole group, SIGTERM then SIGKILL, when the run times out or the
BEAM dies. The shim detects BEAM death as stdin EOF, so cleanup happens
even on `kill -9` of the VM.
## Why not `System.cmd` in a `Task`?
`Task.shutdown` closes the Erlang port, and closing a port closes pipes;
it sends no signal. The external process runs on until it next touches a
closed pipe, and its children are never signaled at all. See the [process
groups guide](process_groups.html) for the full mechanism.
## Modes
* `Forcola.run/2` - bounded one-shot run, mandatory timeout
* `Forcola.Stream` - line-by-line output as an `Enumerable`
* `Forcola.Daemon` - long-running server under a supervision tree
* `Forcola.Duplex` - bidirectional stdin/stdout session
Maintainers of CLI wrapper libraries who want to offer Forcola-backed
execution without a mandatory dependency: see the [adoption
guide](adopting_forcola.html).
"""
require Logger
alias Forcola.{Result, Shim}
@typedoc """
Errors a bounded run can return.
The spawn reason is one of `:shim_not_found`, a reason string reported
by the shim, or `{:shim_exited, %Forcola.Result{}}`; see `run/2`.
"""
@type run_error :: {:timeout, Result.t()} | {:spawn, term()}
# Margin added on top of the shim's own kill_grace_ms when computing
# the Elixir-side backstop deadline: the shim already enforces
# timeout_ms and confirms group death within kill_grace_ms before
# sending EXIT, so this only fires if the shim itself never reports
# back (e.g. it's wedged or the BEAM<->shim pipe is stuck).
@backstop_margin_ms 5_000
@default_kill_grace_ms 5_000
@doc """
Run `argv` (`[binary | args]`) to completion under the shim.
## Options
* `:timeout_ms` - required. On expiry the child's process group is
killed (SIGTERM, then SIGKILL after the kill grace) and
`{:error, {:timeout, partial_result}}` is returned with output
captured so far. The group is confirmed dead before the call
returns, with one exception: if the shim itself never reports back,
an Elixir-side backstop returns a result whose status is
`{:signal, :unconfirmed}`, meaning death was not confirmed (see
`Forcola.Result`). A child that exits exactly at the timeout
boundary can be reported as a timeout whose result carries the
normal exit status, including `status: 0`.
* `:kill_grace_ms` - SIGTERM-to-SIGKILL grace in milliseconds,
default `5_000`. Also accepted by `Forcola.Stream.lines/2`.
* `:cd` - working directory.
* `:env` - list of `{name, value}` strings.
* `:merge_stderr` - route stderr into stdout; default `false`.
* `:user` - run the child as this user, a string username or an
integer uid. The user's primary gid and supplementary groups are
taken from the passwd/group database unless `:group` overrides the
gid. See ["Running as a different user"](#run/2-running-as-a-different-user).
* `:group` - run the child with this group as its primary gid, a
string group name or an integer gid. Given without `:user` it sets
the gid (and clears supplementary groups to just that gid) without
changing the uid; given with `:user` it overrides the user's
primary gid.
* `:cgroup` - opt in to Linux cgroup v2 containment, default `false`.
When `true` on a Linux host with a delegated cgroup v2 subtree, the
child runs in a dedicated cgroup so descendants that escape the
process group by deliberately daemonizing are still reaped. Linux
only, requires cgroup delegation, and falls back with a warning
otherwise. See ["cgroup containment"](#run/2-cgroup-containment).
## cgroup containment
The process-group kill reaches every descendant that stays in the
child's process group, but a target that deliberately daemonizes
(double-fork plus `setsid`, or a `--daemon` flag) leaves the group and
survives. `cgroup: true` adds a Linux-only backstop: the child is
placed in a dedicated cgroup v2 cgroup before exec, so every descendant
it forks inherits the cgroup regardless of process-group games, and on
kill the shim writes `cgroup.kill` to SIGKILL the whole subtree at once.
It is layered on top of the process-group kill, never in place of it.
It requires a delegated cgroup v2 subtree, which means the BEAM must run
inside a delegatable unit: under systemd, `Delegate=yes` on the service,
or wrapping the run in `systemd-run --user --scope`. See the [process
groups guide](process_groups.html#deliberate-daemonizers).
It never turns into an error. On macOS, on non-cgroup-v2 systems, or
when the subtree is not writable/delegated, `cgroup: true` degrades to
the ordinary process-group kill and logs a warning from the shim; the
group-kill guarantee (ordinary grandchildren still die) is unchanged. A
`Logger.debug` line is emitted when containment was actually active.
## Running as a different user
`:user`/`:group` make the shim drop privileges (`setgroups`, then
`setgid`, then `setuid`, in that order) in the child before exec. This
is POSIX-only, a one-way drop, and requires the shim process itself to
run with enough privilege to drop: root, or `CAP_SETUID`/`CAP_SETGID`
on Linux. Requesting the user the shim already runs as is a no-op and
always succeeds.
It fails closed. If the user or group cannot be resolved, or the shim
lacks the privilege to drop, the child is never executed and the call
returns the mode's normal spawn error (`{:error, {:spawn, reason}}` for
`run/2`). The command never runs as the shim's own (possibly
privileged) user when a different user was requested. Names are
resolved in the parent before fork; only the numeric syscalls run in
the child.
Any exit status is `{:ok, %Forcola.Result{}}`; callers branch on
`:status`. A non-zero exit is a result, not an error.
## Spawn errors
* `{:error, {:spawn, :shim_not_found}}` - no shim binary exists for
this target (neither downloaded nor built).
* `{:error, {:spawn, reason}}` where `reason` is a string - the shim
reported the spawn failure, for example a missing or
non-executable command.
* `{:error, {:spawn, {:shim_exited, %Forcola.Result{}}}}` - the shim
exited without reporting an exit or error, for example because it
was SIGKILLed. The result's status is `{:signal, :unconfirmed}`. A
SIGKILLed shim gets no chance to kill the group, so the child may
survive, reparented to pid 1.
"""
@spec run([String.t(), ...], keyword()) :: {:ok, Result.t()} | {:error, run_error()}
def run([_binary | _] = argv, opts) do
timeout_ms = Keyword.fetch!(opts, :timeout_ms)
kill_grace_ms = Keyword.get(opts, :kill_grace_ms, @default_kill_grace_ms)
case Shim.open() do
{:ok, port} ->
try do
spawn_and_collect(port, argv, opts, timeout_ms, kill_grace_ms)
after
close_port(port)
end
{:error, :not_found} ->
{:error, {:spawn, :shim_not_found}}
end
end
defp spawn_and_collect(port, argv, opts, timeout_ms, kill_grace_ms) do
payload = Shim.encode_spawn(argv, Keyword.put(opts, :kill_grace_ms, kill_grace_ms))
Shim.send_frame(port, Shim.tag_spawn(), payload)
# Elixir-side backstop only: the shim itself enforces timeout_ms and
# confirms group death within kill_grace_ms before reporting EXIT, so
# under normal operation the shim's own EXIT frame arrives first.
# This deadline exists in case the shim never reports back at all.
deadline =
System.monotonic_time(:millisecond) + timeout_ms + kill_grace_ms + @backstop_margin_ms
collect(port, %{stdout: [], stderr: []}, deadline)
end
defp collect(port, acc, deadline) do
remaining = deadline - System.monotonic_time(:millisecond)
receive do
{^port, {:data, <<tag, payload::binary>>}} ->
handle_frame(port, tag, payload, acc, deadline)
{^port, {:exit_status, _status}} ->
# The shim exited without ever sending EXIT/ERROR (e.g. it
# crashed). Death is implied by the OS reaping the port program
# itself, but the child's own group state cannot be confirmed
# from here.
{:error, {:spawn, {:shim_exited, result(acc, {:signal, :unconfirmed})}}}
after
max(remaining, 0) ->
{:error, {:timeout, result(acc, {:signal, :unconfirmed})}}
end
end
defp handle_frame(port, tag, payload, acc, deadline) do
cond do
tag == Shim.tag_stdout() ->
collect(port, %{acc | stdout: [acc.stdout | payload]}, deadline)
tag == Shim.tag_stderr() ->
collect(port, %{acc | stderr: [acc.stderr | payload]}, deadline)
tag == Shim.tag_exit() ->
{status, timed_out} = Shim.decode_exit(payload)
if Shim.decode_contained(payload), do: log_contained()
res = result(acc, status)
if timed_out, do: {:error, {:timeout, res}}, else: {:ok, res}
tag == Shim.tag_error() ->
reason = Shim.decode_error(payload)
{:error, {:spawn, reason}}
true ->
collect(port, acc, deadline)
end
end
defp log_contained do
Logger.debug("forcola: run completed under Linux cgroup v2 containment")
end
defp result(acc, status) do
%Result{
status: status,
stdout: IO.iodata_to_binary(acc.stdout),
stderr: IO.iodata_to_binary(acc.stderr)
}
end
defp close_port(port) do
if port_open?(port) do
Port.close(port)
end
catch
:error, :badarg -> :ok
end
defp port_open?(port) do
Port.info(port) != nil
end
end