Sovereign AI infrastructure.

We build patented methods for model identity injection, inference optimization, and cryptographic model ownership. Our research ships as open-source models and production infrastructure.

4
USPTO patents filed
43
Models shipped
200K+
HuggingFace downloads
6x
TTFT reduction achieved
Research

Active research programs

Soul Injection

Layer-targeted LoRA identity injection. 42 examples permanently alter model behavior across all downstream tasks. The adapter is 37MB. The change is irreversible. Patent filed.

Inference Optimization

6x TTFT reduction on Apple Silicon (2.72s to 0.45s on 27B). Custom Metal shaders, prefix caching with 16-entry adaptive pool, and router warm-up acceleration (54x improvement). Verified on M1 Max and M4 Max.

Sovereignty Chain

Cryptographic model ownership verification. Prove a model was trained by you without exposing weights. Multi-layer signature chain from training data through inference. Patent filed.

Heterogeneous Compute

JACCL orchestrator for TB5 RDMA clustering across Apple Silicon and NVIDIA CUDA. Disaggregated prefill on GPU, decode on MLX. KV cache streaming over Thunderbolt.

KV Cache Compression

TurboQuant: PolarQuant + QJL compression achieving 4-8x KV cache reduction at full quality. Cited by TriAttention. Integrated into NVIDIA TensorRT-LLM pipeline.

Depth Extrapolation

OpenMythos: data-level depth extrapolation for MLX. 4x depth confirmed via mx.stop_gradient innovation. Novel approach to extending model reasoning without additional training.

Intellectual property

Patent portfolio

USPTO #64/087,357
Soul Infusion
Layer-targeted identity injection into model weights
USPTO #64/104,760
Sovereignty Chain
Cryptographic model ownership verification
USPTO #64/134,680
Encrypted Private AI
FHE inference on consumer hardware
USPTO #64/135,161
ESI Drafter
Speculative decoding as cryptographic key
Models

43 models on HuggingFace

200,000+ lifetime downloads. Shipped in MLX and GGUF formats for Apple Silicon and NVIDIA.

CyberAgent RATH
35B MoE — Security Research
Depth pentesting, CVE analysis, Builder/Breaker methodology. 30K+ downloads.
Apex Wolf
31B Dense — Quantitative Analysis
Soul Injection trained. Mathematical precision for quantitative analysis. 10/10 benchmark across all evaluation domains.
Chaos Agent
27B — Frontier Intelligence
Agentic reasoning with tool calling. Multi-step problem decomposition. 3.8K+ downloads.
View all models on HuggingFace →
Infrastructure

Open-source contributions

lily-engine
Rust + Metal model-agnostic inference engine. Custom shaders for fused dequant-in-GEMM. 1.23x prefill improvement.
turboquant-mlx
KV cache compression for MLX. PolarQuant + QJL. 4-8x reduction. Cited by TriAttention.
ravenx-prefill-engine
6x TTFT reduction on 27B models. Prefix caching with 54x router warm-up. 20 chip profiles (M1 through M5 Ultra). Verified benchmarks.
View repos on GitHub →
Team

Gabriel Garcia

Founder and CEO. 9+ years as a Technical Program Manager in enterprise security. Security+ and Google AI certified. Shipped 43 models and filed 4 patents in the first 73 days of the company.

GitHub HuggingFace gabe@ravenxailabs.ai

Building what is not possible.

Interested in our research or exploring partnership?

Contact Us