Every traced method call routes through one sink, so its per-call cost has to be nearly invisible. This is a walk through that hot path: what runs per traced call, why it is C, where the time goes, and how failures are contained. Backpressure and invariant handling are decided and shipped: a bug in the tracer degrades tracing, never the host.
native_ext.c: the tracer calls RuntimeAnalysis::Native.emit (or .emit_dict / .roll); the C entry copies under the GVL and push_slot enqueues. The GVL boundary runs through the file — consumer_main drains without it, resolves interned strings through the consumer-owned arena, and writes offline files. A shunt forwards the stream to the hosted graph.Why this is native in the first place
Three forces push the sink out of Ruby and into C:
- Don’t make garbage. A Ruby sink would allocate per call — hashes, strings, arrays — and every one is future GC work in the host’s heap. In C we copy into memory we own and never touch Ruby’s heap again.
- Get the expensive work off the calling thread. JSON encoding and file writes are slow relative to a method call. The design is a single-producer / single-consumer ring buffer: the traced thread (producer) does the cheapest possible thing — copy and enqueue — and a background pthread (consumer) encodes and writes later, on its own time.
- Precise control over threads and memory. We need a real OS thread that runs without Ruby’s
Global VM Lock (GVL), and hand-managed buffers whose lifetime we control across
fork.
There is one sink per (pid, tid) — its own ring, consumer thread, output file, and string arena,
so streams never interfere. Ruby opens a sink and gets back an integer handle; every emit/roll/
close addresses that handle.
The consumer never touches a Ruby object
The consumer thread never touches a Ruby
VALUE.
The consumer runs on a raw pthread with no GVL. If it ever dereferenced a Ruby object, GC — especially
compaction, which moves objects — could pull the ground out from under it. So by the time anything
crosses into the ring it is a native copy: the producer, still holding the GVL, does all the
Ruby→C copying; the consumer only ever reads plain bytes and int64s. This invariant is why releasing
the GVL under backpressure is even thinkable, and every change has to preserve it.
Step by step: one traced call
The functions below are the ones a single traced call actually touches, verbatim; the setup and
teardown around them (open, close, fork handling, the handle registry) are elided.
static VALUE native_emit(VALUE self, VALUE handle, VALUE record) {
sink *s = sink_for(handle);
Check_Type(record, T_ARRAY);
if (RARRAY_LEN(record) != REC_FIELDS) {
rb_raise(rb_eArgError, "runtime_analysis native record must have %d fields", REC_FIELDS);
}
slot sl;
memset(&sl, 0, sizeof(sl));
sl.kind = SLOT_RECORD;
for (int i = 0; i < REC_FIELDS; i++) {
sl.rec[i] = NUM2LL(RARRAY_AREF(record, i));
}
push_slot(s, &sl);
return Qnil;
}
1 · native_emit — the whole cost of a call
The Ruby layer has already interned every string to an integer, so a call arrives as a fixed-width
14-integer record. Validate it, copy each field with NUM2LL into a stack slot — no allocation, no
Ruby object retained — and enqueue. The string case is the sibling native_emit_dict, which mallocs
a native copy of the bytes once per string. All of this runs under the GVL, and it is the producer’s
entire contribution to a traced call.
static void push_slot(sink *s, slot *sl) {
pthread_mutex_lock(&s->mu);
if (s->disabled || s->count == s->cap) {
s->dropped++;
pthread_mutex_unlock(&s->mu);
free(sl->dict_str.ptr); /* NULL for records; the owned copy for dict/roll, so a drop can't leak */
return;
}
s->slots[s->head] = *sl;
s->head = (s->head + 1) % s->cap;
s->count++;
pthread_cond_signal(&s->not_empty);
pthread_mutex_unlock(&s->mu);
}
2 · push_slot — into the ring, never blocking
Lossy by design. The producer holds the GVL, so a blocked producer would freeze every Ruby thread.
When the ring is full, or the sink has been disabled, drop the slot, free any owned copy so a drop
can’t leak, and bump a dropped counter surfaced to Ruby. Otherwise drop it in, bump head mod cap,
and signal the consumer. The ring is guarded by s->mu, a pthread mutex, not the GVL.
static void *consumer_main(void *arg) {
sink *s = (sink *)arg;
for (;;) {
pthread_mutex_lock(&s->mu);
while (s->count == 0 && s->running && !s->disabled) {
pthread_cond_wait(&s->not_empty, &s->mu);
}
if (s->disabled || (s->count == 0 && !s->running)) {
pthread_mutex_unlock(&s->mu);
break;
}
if (!(s->count > 0 && s->tail < s->cap)) {
fprintf(stderr, "consumer invariant failed (%s:%d)\n", __FILE__, __LINE__);
#ifdef RA_ABORT_ON_INVARIANT
abort();
#else
s->disabled = 1;
pthread_mutex_unlock(&s->mu);
break;
#endif
}
slot sl = s->slots[s->tail];
s->tail = (s->tail + 1) % s->cap;
s->count--;
pthread_cond_signal(&s->not_full);
pthread_mutex_unlock(&s->mu);
consume_slot(s, &sl); /* a guard here may disable us; the next loop exits */
}
if (s->file) fflush(s->file);
return NULL;
}
3 · consumer_main — the background thread
Runs on a raw pthread with no GVL, so it may never touch a Ruby VALUE — it only reads native
copies. Wait for work, pop one slot under the lock, release the lock immediately so the producer isn’t
blocked during slow work, then encode and write off the hot path. The invariant check disables just
this one sink on a violation — abort() only under the CI/sanitizer flag — and the loop exits. The
host lives on.
static void consume_slot(sink *s, slot *sl) {
switch (sl->kind) {
case SLOT_DICT:
RA_GUARD(s, s->file, );
consume_dict(s, sl);
s->bytes_written = ftello(s->file);
note_io_error(s);
break;
case SLOT_RECORD:
RA_GUARD(s, s->file, );
consume_record(s, sl);
s->bytes_written = ftello(s->file);
note_io_error(s);
break;
case SLOT_ROLL:
consume_roll(s, sl);
break;
default:
RA_GUARD(s, 0 && "unknown slot kind", );
}
}
4 · consume_slot — dispatch
A record or dict writes a row and updates bytes_written, the segment size Ruby polls to decide when
to rotate; a roll swaps the file. Every RA_GUARD sits before the dangerous access, so a missing
file handle or an unknown kind disables the sink rather than corrupting anything, and a write error is
reported once through note_io_error.
static void consume_record(sink *s, slot *sl) {
int64_t *r = sl->rec;
size_t uuid_id = (size_t)r[F_UUID];
const char *uuid = "";
long uuid_len = 0;
if (uuid_id < s->arena_cap && s->arena[uuid_id].ptr) {
uuid = s->arena[uuid_id].ptr;
uuid_len = s->arena[uuid_id].len;
}
fputs("[1,", s->file);
write_json_string(s->file, uuid, uuid_len);
for (int i = 1; i < REC_FIELDS; i++) {
fputc(',', s->file);
if (i == F_C_CLASS_METHOD || i == F_M_CLASS_METHOD) {
write_bool(s->file, r[i]);
} else {
fprintf(s->file, "%lld", (long long)r[i]);
}
}
fputs("]\n", s->file);
}
5 · consume_record — the row, encoded
Resolve the uuid id to its string through the arena (id → dstr, kept across rotations), then write
[1, "uuid", …] — the two class-method fields as true/false, the rest as integers, byte-for-byte
the offline dictionary-encoded row. write_json_string escapes control characters by hand; bytes ≥
0x20 pass through verbatim, so a non-UTF-8 string yields technically-invalid JSON, deliberately.
static void consume_roll(sink *s, slot *sl) {
char *newpath = sl->dict_str.ptr;
if (newpath) {
FILE *nf = fopen(newpath, "a");
if (nf) {
if (s->file) { fflush(s->file); fclose(s->file); }
s->file = nf;
s->bytes_written = 0;
} else {
fprintf(stderr, "could not open segment %s: %s\n", newpath, strerror(errno));
}
free(newpath);
}
}
6 · consume_roll — rotation as an ordered event
The roll is a slot in the same FIFO, so the segment boundary is exact even though Native.roll
returned to Ruby long before the file was swapped. Open the new file before closing the old one, so a
failed fopen leaves the current stream intact. The arena is kept: a record in the new segment may
resolve a uuid interned in an earlier one.
Where the time goes
- File I/O in the consumer — the dominant cost, and the reason the ring exists.
- JSON encoding — per-byte escaping,
fprintfper field. malloc+memcpyper new string inemit_dict— only on first sighting; interning amortizes it.- Mutex hand-off between producer and consumer — cheap, uncontended in the common case.
The producer touches none of the first two.
The safety bits
- GC safety — the copy-out invariant. Nothing in the ring or arena is a Ruby object, so compaction can move the world and the consumer does not care.
- The ring mutex, not the GVL, guards the ring — producer and consumer are correct independent of Ruby’s scheduler.
RA_GUARDchecks things that must be impossible by construction — ring indices in range, arena bounds, valid slot kinds. Each guard sits before the dangerous access, so a violation means we are about to step outside our own buffers, not that we already have. So we refuse the access and disable that one sink — log to stderr, stop its consumer — leaving the host process alive; the Ruby layer opens a fresh sink on its next call tree. Test and sanitizer builds defineRA_ABORT_ON_INVARIANTtoabort()loudly instead, so a real bug is caught in CI. A tracer must never take down its host, but it must scream during development.forkhandling — the child abandons every inherited sink wholesale. It deliberately does not flush or close inheritedFILE*s: their buffers hold the parent’s unwritten bytes, and flushing our duplicate fd would write them twice. The child re-opens its own sinks.- The GVL serializes the registry —
open/close/resetrun under the GVL and never release it while mutating the handle table, so it needs no extra lock.
How failures are contained
The governing principle: a bug in the tracer — a full ring, a failed allocation, or a broken invariant — degrades tracing, never the host. The producer holds the GVL, so anything that blocks or crashes it hurts the whole VM; every failure mode below is handled without doing either.
- Backpressure → lossy drop, never block.
native_emit/emit_dict/rollhold the GVL, so a blocked producer would freeze every Ruby thread. Instead, when the ring is fullpush_slotdrops the slot, frees any owned copy, and bumps adroppedcounter surfaced to Ruby (Native.dropped). Offline, a dropped record is one missing call edge — a gap in that tree’s step sequence; a dropped dict leaves later records resolving to an empty label, degraded but never mis-attributed. Under steady pressure the dictionary is already warm, so drops are almost all records. The stream is honestly incomplete, counted, never silently wrong. - Invariant violation → disable the sink, not the host. The
RA_GUARDchecks above fire before the dangerous access, so we disable that one sink and the Ruby layer opens a fresh one at its next call tree. Blast radius is one call tree, not the process.abort()only under the CI/sanitizer flag. - Allocation failure → drop or raise, never segfault. Every allocation is NULL-checked. On the hot
path an OOM drops and counts rather than raising into the traced method; at setup it raises a clean
Ruby exception, unwinding the half-built sink first. A negative
capacityis rejected withArgErrorbefore it can wrap to a hugesize_t. - I/O failure → visible, not silent. A failed
fopeninconsume_rollleaves the current segment intact and logs to stderr; consumer write errors are reported once viaferror, so the host can learn the stream went incomplete.
Self-healing ties it together: because a disabled sink’s guards fire before any corrupting write, its
buffers are still valid, so the Ruby layer can close it and open a fresh segment — recovery, not just
survival.