What's been cooking — July 2026
7 merged PRs across 3 repos
July was mostly a CUDA month: a lot of time spent staring at nsys traces, chasing non-determinism in atomics, and discovering that the inference pipeline was doing a PIL round-trip nobody asked for. Falcata soaked up most of the attention, with smaller but satisfying surgical fixes landing in NVIDIA's stack.
MechaFauna-ai/Falcata
Five PRs, all with teeth. The headline performance win is Falcata#27. nsys showed per-split CUDA training was host-sync bound, with cudaDeviceSynchronize eating about 45% of wall time against 21% for all the GPU kernels combined. The PR swaps the two device syncs between histogram construction and FindBestSplits for event-based stream ordering, so the two child leaves run concurrently with no host stall: +8.6% mean end-to-end training speed on a 500k×100, 255-leaf fit (interleaved A/B, n=30, Welch's t significant). Falcata#22 fixes a genuinely nasty one in CUDADataPartition::CalcBlockDim: rounding per-block data counts up to a power of two made the grid non-monotonic in leaf size, so a smaller leaf could demand more blocks than the full dataset, and the block-offset buffers, sized once for the full dataset, got written past the end. compute-sanitizer caught it; on most allocators it had been silently corrupting adjacent device memory while predictions stayed bit-identical, which is why nobody noticed.
On the correctness side, Falcata#14 makes the LambdaRank gradient kernel bit-deterministic when deterministic=true, replacing the atomicAdd_block scatter whose FP non-associativity meant two runs could produce two different trees. It's opt-in for a reason: the deterministic kernels cost anywhere from +12% to +376% depending on items per query. Rounding out the month, Falcata#40 finally makes device_type=cuda build and run on Windows under MSVC (single-GPU only, since NCCL still doesn't ship for Windows), and Falcata#39 turns on warnings-as-errors, documents the NCCL build, and fixes the four latent issues the new -Werror surfaced, plus a C++ test target that hadn't compiled since the rename.
NVIDIA/Isaac-GR00T
One PR but a good one: Isaac-GR00T#727 deletes the PIL and tensor round-trips from the inference image path. _apply_vlm_processing was helpfully converting every CHW uint8 frame to PIL before handing it to Qwen2VLImageProcessorFast, which accepts the tensor directly and produces identical output. Nine lines changed in the source, and preprocessing per control step drops 63–70% (1.84 → 0.69 ms for one camera, 5.93 → 1.76 ms for three), with a parity suite asserting bitwise-identical processor outputs against the old PIL path.
NVIDIA/cuvs
A tidy micro-optimisation in cuvs#2128: in the filtered brute-force CSR path in knn_brute_force.cuh, the sparse ≥0.9 branch was unconditionally building a per-nonzero rows array via csr_to_coo, even when the inner-product kernel didn't need it. Only the L2 and cosine epilogues consume it, so for inner-product searches the allocation and the extra kernel launch were dead work on every call. Now they only happen when needed. Eight lines each way, no behaviour change.
That's the month — fewer features, more "why is the GPU waiting on the CPU". See you in August.