Pablo Soage

Systems · Reverse Engineering · FPGA

Computer engineer working close to the hardware: undocumented protocols, real-time signal processing in programmable logic, and the kind of problem that is only solved by measuring it. Currently a core software developer at Ágata Technology, and building toward embedded and defence systems.

public repos
stars
languages
last push

Selected work

Repository metadata below — language, stars, last push — is read live from the GitHub API, so it is never out of date with the code.

Upstream

Other people's projects, made better.

codebase-memory-mcp #1764

Every idle client burned about 0.7 of a CPU core from one release onward, and the cost multiplied by the number of open sessions. Bisected against the last good version as a control and reproduced by starting the binary from a shell and sending it nothing at all. Fixed upstream; on closing, the maintainer wrote that “the diagnosis in this thread did the hard part”.

DiagnosisProfilingWindows

nvidia-pstated #10

Datacentre GPUs such as the P100 and V100 expose a single performance state, so the tool had nothing to switch them to. This adds a clock-based fallback: where setting a P-state fails, read each GPU's lowest supported core and memory clocks and idle there instead. Open.

CNVMLTesla V100

Languages, measured

Bytes of source across the projects listed above, markup and build glue excluded.

reading the API…

Read per repository from the API, so this is real source rather than repository size — counting the latter would put a folder of datasets above every line of C here. LaTeX, HTML and build files are left out for the same reason: the thesis alone carries half a megabyte of typesetting. Even then, bytes are not where the work went. A Flex/Bison parser is 60 KB and an RTL design is smaller still, and the reverse-engineering work lives in private repositories that this cannot see.

The bench

What the work above actually runs on.

Compute

chassis
Gigabyte T181-G20, bare metal
cpu
2 × Xeon Platinum 8171M, 52 cores
gpu
4 × Tesla V100 SXM2, NVLink
runs
local LLMs (vLLM, Open-WebUI), Docker swarm
driven by
redfishctl, over Redfish

Programmable logic

board
AMD Kria KV260, K26 SOM
device
Zynq UltraScale+ ZU5EV
carries
scanner64, 64-channel DDC bank
measured
6 400 M channel-samples/s at 0.535 W

Vehicle

interface
Scanmatik SM3, J2534 over Wi-Fi
bus
GMLAN, 500 kbps, ISO-TP
tooling
opendash, protocol recovered from captures
verified
against a running engine