Inference
Find the capacity limits, saturation points and routing decisions that determine serving cost, latency and reliability at scale.
NCA Research
Inference · Optimization · Retrieval
AI infrastructure / model systems
NCA helps organizations make better infrastructure and model-system decisions. Independent measurement turns uncertain architecture choices into clear trade-offs across performance, cost and production risk.
Where NCA creates leverage
Find the capacity limits, saturation points and routing decisions that determine serving cost, latency and reliability at scale.
Turn hardware behaviour into higher utilisation and better unit economics across NVIDIA and AMD inference stacks.
Build compact multimodal retrieval systems that preserve quality while reducing index, memory and downstream context costs.
Visual retrieval / SOTA models
Built and evaluated on Vultr Cloud, the three-tier VultronRetriever family pairs state-of-the-art visual retrieval with production-efficient footprints. Prime ranks #1 on ViDoRe V3 with an 8–16× smaller index than its peers, while the 0.8B Flash tier brings the same 320-dimensional design to a ≈1.6 GB BF16 footprint.
Selected research
Distributed inference / 2026
Shows how shared prefill and decode capacity degrades at saturation, and how adaptive routing can recover a better operating point.
3.1× lower contention penalty in the strongest saturated case.
GPU performance / 2026
A practical guide to the configuration choices that change throughput, stability and deployability on AMD Instinct MI325X.
235B–1T models tested across 17,406 requests.
Model architecture / 2026
Combines retrieval and generation in one vision-language model, removing the cost and complexity of duplicated model stacks.
62.7% lower peak GPU memory at the 4B scale.
Document retrieval / 2025
Returns the relevant regions of a document instead of sending entire pages through the downstream vision stack.
52.3% fewer context tokens than full-page image input.
Operator notes / Athrael.net
Practical analysis of what changes capacity, cost and model-system performance—and what only looks good in a benchmark.
The director's cut of the B200 experiments: from three coupled games to the 270-line controller that reacts at saturation.
Why a late-interaction student should not inherit hard negatives chosen by a single-vector cosine scorer, with the measurements to show it.
A full optimisation trail showing when HPO still transfers and when leaderboard tuning starts reshuffling gains between tasks.
Independent judgement
NCA combines independent research with more than 15 years of building production systems. For leaders deciding where to place capital and engineering effort, that means a clear view of what scales, what saturates and what is ready to deploy.
Work spans NVIDIA CUDA and AMD ROCm inference, vLLM, NVIDIA Dynamo, multimodal late interaction and billion-scale vector retrieval. The output is decision-grade evidence: measured systems, deployable models and technical recommendations that survive real hardware and real traffic.
Research · Engineering · Advisory
NCA works with technical and executive leaders on inference strategy, system optimization and retrieval—especially where capital efficiency, performance and production constraints meet.
Start a conversation