Security

colibri: heap overflow in the safetensors loader

Heap out-of-bounds write in c/st.h safetensors loader: unvalidated tensor metadata overflows heap buffer on model load (GHSA-4gw4-j89j-4c8r)

PUBLISHED SEVERITY HIGHSTATUS PatchedGHSA-4gw4-j89j-4c8r
Patched

Fixed in colibri 1.3.0 (2026-07-29) per the advisory. The loader now checks each tensor's byte span against the file, its shape against its byte span, and the copy length against the destination buffer. Track the fix at JustVugg/colibri#413.

Product
colibri (C inference engine) in JustVugg/colibri
Affected versions
colibri before 1.3.0, per the advisory. Reviewed at 62419af (2026-07-14). The main checks first shipped in 1.1.0 (2026-07-22); the advisory names 1.3.0 (2026-07-29) as the first patched release.
Severity
HIGH
Status
Patched
Weaknesses
  • CWE-20Improper Input Validation
  • CWE-787Out-of-bounds Write
  • CWE-1284Improper Validation of Specified Quantity in Input
As filed on the advisory.
GitHub advisory
GHSA-4gw4-j89j-4c8r
Published by the maintainer on 2026-08-05
CVE
Pending (requested from MITRE)
Published
Credit
Finder: Aaron Elijah Mars of Aeon. Tool: Aeon (https://www.aeon.fun).

What users should do now

  1. Upgrade colibri to 1.3.0 or later. The latest release is v2.0.0.
  2. Load model snapshots only from sources you trust, especially on builds older than 1.3.0.
  3. If a fixed build refuses a model with a data_offsets ... out of file bounds, shape/bytes mismatch or exceeds destination capacity error, treat that file as hostile or corrupt and do not work around the check.

Summary

colibri runs large mixture-of-experts models from local safetensors files, and its README points users to third-party converted snapshots on model mirrors. The loader in c/st.h read each tensor's data_offsets, shape and dtype from the file header and stored them as given. Later reads used those values to decide how many bytes to copy into a buffer that the engine had sized separately. When the header disagreed with that size, the copy ran past the end of the heap buffer. The engine already treated config.json as untrusted input and range-checked it, but the safetensors header ships in the same bundle and had no such checks.

Affected versions

colibri before 1.3.0, as stated on the advisory, when it loads a model directory whose safetensors files the attacker wrote or modified. We reviewed commit 62419af (2026-07-14). The checks landed in stages: the offset and header checks shipped in 1.0.0, the shape and destination-size checks in 1.1.0, and the slice bounds fix in 1.2.0. The maintainer marked 1.3.0 as the first patched release.

Impact

  • Heap out-of-bounds write on model load. A crafted file can make the loader write data it controls past the end of a heap buffer. This happens while weights load, before any prompt is processed. Heap corruption of this kind can crash the engine and may allow code execution as the user running colibri.
  • Out-of-bounds read. The same mismatch between shape and byte span made the BF16 and F16 conversion loops read past the end of the temporary buffer that holds the raw tensor data.
  • Crashes from malformed headers. A header that left out dtype, data_offsets or shape, or used the wrong type for them, was dereferenced without a check. The header length from the file was also used for an allocation without a bound.

The attacker needs the victim to load a model file they supplied, for example a converted snapshot published on a model mirror. No network access to the victim's machine is needed.

Affected code

Permalinks at 62419af, the reviewed commit. At that point the GLM engine lived in c/glm.c; it was later split into c/colibri.c and header modules.

  • st_init() reads data_offsets and shape and stores nbytes = b0 - a0 and numel (the product of the shape) as given. It does not check that the offsets are ordered and inside the file, that numel matches nbytes for the dtype, or that the shape product does not overflow. It also uses the fields without NULL or type checks: c/st.h#L121-L141
  • st_read_f32() copies nbytes into the caller's buffer for F32 tensors, and writes numel floats for BF16 and F16 tensors. Neither is bounded by the size of the caller's buffer: c/st.h#L183-L198
  • st_read_raw() and st_read_slice_f32() also take their lengths from the header and have no destination bound: c/st.h#L209-L231
  • Callers size the destination separately from the copy length. qt_from_disk() allocates the per-row scale buffer from the config (O floats) and then calls st_read_f32() on the .qs scale tensor; ld() allocates from the declared shape while the F32 copy uses the byte span: c/glm.c#L1059-L1088
  • For comparison, config.json was already range-checked because it comes from untrusted mirrors: c/glm.c#L1038-L1051

Fix

  • 4268e00 bounds the header length by the file size, checks that dtype, data_offsets and shape are present and have the right type, and rejects offsets that are negative, out of order or past the end of the file. Shipped in 1.0.0.
  • #413 requires numel * element size == nbytes before any copy, guards each multiply in the shape product against overflow, and makes the quantized-weight loader refuse weight and scale sizes that do not match the config shape. Shipped in 1.1.0.
  • #368 adds st_read_f32_cap(), which refuses to write more elements than the destination holds, and uses it wherever the buffer is sized from the config. It also adds a matching size check on quantized tensors. Shipped in 1.1.0.
  • #618 keeps st_read_slice_f32() reads inside the tensor without integer overflow, and rejects a negative element count in the capped reader. Shipped in 1.2.0.
  • The advisory lists 1.3.0 as the first patched release. All of these checks are still present on main at bf24429 (v2.0.0).

Detection (for defenders)

On affected builds, look for colibri crashing or aborting with heap corruption while it loads a model, especially a model fetched from a third-party mirror. On fixed builds, the loader exits with a clear error naming the tensor (out of file bounds, shape/bytes mismatch, exceeds destination capacity or malformed dtype/data_offsets/shape); any such error on a downloaded model means the file should not be trusted. A working exploit is withheld.

Timeline

  1. Aeon reviews colibri at 62419af and finds the unchecked tensor metadata.
  2. Maintainer commits 4268e00 (header length, missing-field and offset checks).
  3. #413 merged (shape and byte-span cross-check, shape overflow guard).
  4. Reported privately via GitHub PVR (GHSA-4gw4-j89j-4c8r), with a suggested fix.
  5. #368 merged (capped reads into config-sized buffers).
  6. colibri 1.1.0 released with #413 and #368.
  7. #618 merged (slice bounds and negative count checks); shipped in 1.2.0 on 2026-07-28.
  8. colibri 1.3.0 released, the first patched release per the advisory.
  9. Maintainer publishes the advisory alongside colibri 1.5.0, noting the issue was already closed by #413 and re-verified.
  10. CVE ID requested from MITRE.
  11. Maintainer asked to request a CVE through GitHub (#2002); public write-up.

Credit

Finder: Aaron Elijah Mars of Aeon. Tool: Aeon.

References