We build patented methods for model identity injection, inference optimization, and cryptographic model ownership. Our research ships as open-source models and production infrastructure.
Layer-targeted LoRA identity injection. 42 examples permanently alter model behavior across all downstream tasks. The adapter is 37MB. The change is irreversible. Patent filed.
6x TTFT reduction on Apple Silicon (2.72s to 0.45s on 27B). Custom Metal shaders, prefix caching with 16-entry adaptive pool, and router warm-up acceleration (54x improvement). Verified on M1 Max and M4 Max.
Cryptographic model ownership verification. Prove a model was trained by you without exposing weights. Multi-layer signature chain from training data through inference. Patent filed.
JACCL orchestrator for TB5 RDMA clustering across Apple Silicon and NVIDIA CUDA. Disaggregated prefill on GPU, decode on MLX. KV cache streaming over Thunderbolt.
TurboQuant: PolarQuant + QJL compression achieving 4-8x KV cache reduction at full quality. Cited by TriAttention. Integrated into NVIDIA TensorRT-LLM pipeline.
OpenMythos: data-level depth extrapolation for MLX. 4x depth confirmed via mx.stop_gradient innovation. Novel approach to extending model reasoning without additional training.
200,000+ lifetime downloads. Shipped in MLX and GGUF formats for Apple Silicon and NVIDIA.
Founder and CEO. 9+ years as a Technical Program Manager in enterprise security. Security+ and Google AI certified. Shipped 43 models and filed 4 patents in the first 73 days of the company.