RightNow Research Lab
Enabling Model-Hardware Co-Design at Scale

Abstract
We build tools that close the gap between AI models and the hardware they run on.
NVIDIA Inception member · 6 papers on arXiv

Endorsement
Director of Accelerated Computing · NVIDIA
Products

Recorded on runinfra.ai on 2026-10-09: the home page, Open models, built for agents, with a model's speed, price and cache-hit panel, then that model's API call in Python and the Anthropic SDK.
Hosted open models
RunInfra
One API key for the OpenAI and Anthropic SDKs. Billed per token from your balance.
Explore RunInfra (opens in a new tab)Research

2606.09682cs.LG
AutoMegaKernel: A Statically-Checked Agent Harness for Self-Retargeting Megakernel Synthesis
Jaber Jaber*, Osama Jaber · 2026

2604.02051cs.LG
Ouroboros: Dynamic Weight Generation for Recursive Transformers via Input-Conditioned LoRA Modulation
Jaber Jaber*, Osama Jaber · 2026

2603.29090cs.LG
HCLSM: Hierarchical Causal Latent State Machines for Object-Centric World Modeling
Jaber Jaber*, Osama Jaber · 2026

2603.21365cs.LG
TIDE: Token-Informed Depth Execution for Per-Token Early Exit in LLM Inference
Jaber Jaber*, Osama Jaber · 2026
Agent OS
OpenFang
Low-level agent execution with direct OS and GPU access.
RightNow-AI/openfangOpen repo
OpenFang: GitHub Stars
18,209
GitHub API count: 2026-09-21; chart illustrative
Edge inference
PicoLM
Small inference runtime for memory-constrained edge devices.
RightNow-AI/picolmOpen repo
PicoLM: GitHub Stars
2,193
GitHub API count: 2026-09-21; chart illustrative
Kernel search
AutoKernel
Profiles PyTorch models, tests Triton variants, and follows measured bottlenecks.
RightNow-AI/autokernelOpen repo
AutoKernel: GitHub Stars
1,564
GitHub API count: 2026-09-21; chart illustrative




