Three Ways Out of the Token Loop: HyperZip, Jev and Janus
Published October 2, 2026 · 2193 words · 11 min read
Every GatiFlow report from 11:09 UTC on September 21 through September 30 ranked the same kind of repository first among GitHub projects by star growth: something built on or around Jev, the decision model TypeSafe AI opened to early access on September 15. On September 30, GatiFlow's arXiv collector flagged HyperZip, which uses diffusion language models to compress data faster, and Janus, a system that keeps the memory of long agent sessions on SSDs, as newly active papers. One signal is loud and one is quiet, and they look unrelated. They are the same story.
A language model normally writes one token per forward pass and carries a growing memory of the session from step to step, and each of these projects is trying to shrink one of those two costs. There are three routes out of the token loop: produce many tokens per pass, skip generation when a decision is all the software needs, or make the session's memory cheaper to keep. Each route already has something running in production or in open source, which is what makes this a Saturday subject rather than a Wednesday headline.
Start with what the data does and does not show. In Wednesday's 17:01 UTC report, GatiFlow's arXiv collector marked five research papers as newly active, each seen by that one collector and nowhere else. HyperZip was submitted on September 28; Janus, BRIDGE and OmniRoute on September 29; the fifth, Exemplar2VQA, is a generator of visual question-answering data. Three of the five are about cost: HyperZip attacks decoding throughput, Janus the price of reloading session memory, and OmniRoute, a training-free method for compressing audio and video tokens, the prefill cost of omnimodal models.
BRIDGE is not a cost paper, and when the 21:33 UTC report described the batch as "a coordinated turn in research attention," it claimed more than one day of one collector can show. The 17:01 report had it right when it called the cluster "suggestive but not yet the multi-day pattern." What makes it a pattern is the evidence around the papers, starting with HyperZip itself.
HyperZip, by Thai Nguyen, Khang Tran and NhatHai Phan of the New Jersey Institute of Technology, works on a narrow problem that carries a general lesson: using a language model as a lossless compressor. A model that predicts the next symbol well needs fewer bits to encode it, which is why language models have shown strong potential as compressors. The authors write that such compressors are held back by the cost and low throughput of autoregressive decoding.
HyperZip replaces the autoregressive model with a diffusion language model that uses multi-token prediction, so each pass handles several tokens. The authors name the catch themselves: as decoding throughput rises, the compression rate worsens, a metric where lower means smaller output. Their remedy, modeled on Text-to-LoRA, is a hypernetwork that reads an embedding of each document and generates a LoRA update for the frozen diffusion model, so every document gets its own adapted model without a fine-tuning run. Only the small context vector travels with the compressed bits, and personalization moves from a training job to an inference step.
The paper's tables put numbers on the trade. Its appendix states the rule: "All parallel decoding methods improve throughput over the autoregressive baseline, but generally produce larger compressed representations." HyperZip's answer is to claw back size without giving up speed. On documents grouped by topic, it shrinks programming text to 16.3 percent of its original size, against 18.0 percent for the same diffusion model without the hypernetwork and 17.7 percent after conventional fine-tuning, while running at 138.9 tokens per second against 141.0 and 140.2. On the other three topics the gain over the plain diffusion model is 0.4 to 0.5 points. It is an early preprint: four pages of main text, no error bars, no limitations section, and code promised for after acceptance.
It is not a lone idea. On August 4, Angelo Nardone of the University of Pisa and Paolo Ferragina of Scuola Superiore Sant'Anna posted Diffuse to Compress, which proposed diffusion language models for lossless text compression for the same throughput reason and tested the approach on the enwik8 benchmark. Two groups eight weeks apart make a research line, not yet a field.
The commercial anchor is stronger. On September 8, Inception launched Mercury 2.5, a diffusion model that generates tokens in parallel, and said it runs at over 1,100 tokens per second in production, with enterprise usage reaching thousands of developers and dozens of enterprises since Mercury 2. Those are vendor figures. Artificial Analysis, which measures the API as customers call it and with its own method, recorded a little over six hundred tokens per second at the end of September. Either way, the decoding approach HyperZip depends on is something companies can buy, not only a lab method.
The loud signal points the same way, which is easy to miss when star counts are read as hype. TypeSafe calls Jev a System One model: rather than writing text, it returns typed decisions, such as a classification, a route or a score, in a single parallel pass, priced at $0.042 per million input tokens with output free. Three days after the launch, GatiFlow's GitHub collector logged the first repositories built on Jev or on its idea.
The projects show where developers want that trade. In jev-ultrafast, a browser agent from the Browser Use team, Jev picks each operation and the page element it acts on. A Claude Code plugin, tamaratran/fast-jev-compaction, has Jev score every tool call and result in a long session, then drops or truncates the stale ones instead of summarizing.
Others rebuild the model in the open. NandhaKishorM/laya, a "non-autoregressive System 1 decision engine" built on ModernBERT-class encoders, entered GatiFlow's data on September 20 and had 29,229 stars by September 30. A runtime for it on Apple silicon, mizorewww/laya-mlx, advertises decisions in 7 to 14 milliseconds on an M3 Max with no text generation. Jared Palmer, VP of engineering at Cognition, published jaredpalmer/kev, a family of Jev-like decision models on Qwen that users train and run themselves; it went from 6,635 to 8,054 stars after GatiFlow first logged it on September 24. None of these is an agent orchestration framework. They pull routine agent decisions, and the upkeep of agent context, out of the generation loop.
The third route is the session's memory. Agentic sessions alternate between inference and tool calls, and each round adds history whose key-value cache must be either kept or recomputed. Janus, posted on September 29 by Wenhao He and seven co-authors, targets models with sparse attention, which reads only part of that history, and stores the cache on SSDs because they cost less than CPU memory.
The obstacle is that the model only decides which blocks it needs in the middle of computing, which puts slow disk reads on the critical path. Janus borrows the model's own selection module and runs it on intermediate values computed earlier, predicting those reads so they overlap with computation, then fetches any misses before attention runs so outputs do not change. Across three models and three agentic traces, the authors report maximum time-to-first-token speedups between 1.57 and 3.69 times over existing systems, and average speedups of 1.22 to 1.85 times, with decode efficiency preserved.
Janus lands in a market that already exists. LMCache, presented at MLSys 2026, moves KV caches out of GPU memory and shares them across vLLM and SGLang engines. The open-source llm-d scheduler subscribes to cache events from vLLM, SGLang and TensorRT-LLM and routes each request to where its blocks already live. DDN announced KV cache acceleration integrated with NVIDIA Dynamo on March 16 and is hiring a senior software engineering manager for its KV Cache Platform. Even the name is crowded: SOSP, which ended in Prague on October 2, accepted an unrelated Janus for multi-LLM serving, next to a paper on foundation-model serving in Amazon Bedrock.
BRIDGE needs its own line because GatiFlow's reports grouped it with the cost papers. It is about accuracy. Quan Xiao, Tianyi Chen and co-authors argue that most agentic reinforcement learning methods blame the language model for failures that begin with poor retrieval, so they train the retriever and the policy together as a bilevel problem, adapting the retriever first. With 3B and 7B backbones they report multi-hop gains of 9.6 and 3.4 exact-match points over the strongest baseline; the memory efficiency the paper advertises belongs to its training method, not to serving. For teams building search agents the lesson is about where errors come from, not about what tokens cost.
Read separately, the star charts suggest developer fashion and the papers a quiet sideshow. Read together, they show one economic pressure in three places at once: developers rebuilding a decision model in the open within days of its launch, researchers posting compression and caching papers, and vendors selling diffusion decoding and KV cache tiers.
Where the market can over-rotate is in reading attention as adoption. Week-over-week star growth for laya has more than halved since Monday's 22:40 UTC report. Jev is less than three weeks old, laya's claim to beat it on accuracy and latency is self-reported, and HyperZip and Janus are single papers seen by one collector. The hiring layer is the thinnest. Of nearly two thousand distinct job postings that surfaced in GatiFlow's September reports, seven named inference or model serving in the title, among them Crusoe's search for a senior engineering manager in AI inference on September 30, and none asked for KV cache or compression skills. DDN is staffing for this; the broader job market GatiFlow tracks is not yet.
If you are building agents or running high-volume LLM pipelines, the conversation to have with your team in the next two to four weeks is an audit of where your tokens go. First, list every model call whose output is a decision rather than prose, such as routing, classification, yes or no checks, scores and guardrails, and test a single-pass decision model against it on your own labeled sample, hosted like Jev or open like laya and kev, before trusting anyone's benchmark.
Second, for long agent sessions, measure how much time-to-first-token goes to rebuilding context, and check what your serving engine with LMCache or llm-d's cache-aware routing already recovers before you buy memory; Janus is evidence that SSD tiers can work when reads are predicted early, not yet a product. Third, for latency-bound generation, benchmark a diffusion model such as Mercury 2.5 on your own traffic, since the vendor's figure and the independent measurement are far apart. HyperZip is worth a read for teams that store large text datasets, not yet a deployment decision. Each check is cheap now; the architecture it informs gets expensive to change once a serving stack hardens around generating every answer one token at a time.
Forward catalysts are thin. No conference, release window or investor event in the next 7 to 14 days could be confirmed as tied to HyperZip, Janus or the Jev-style models. The nearest relevant date falls just outside that window: PyTorch Conference in San Jose on October 20 and 21, run by the foundation that hosts vLLM, one of the engines LMCache and llm-d work with. Until then, the signals to watch are in GatiFlow's own feed: whether the Jev-style repositories turn stars into forks, issues and package downloads, and whether HyperZip or Janus is picked up by a second collector.
The next round of gains in AI economics may come less from smarter tokens than from fewer, cheaper ones.
Sources:
- [2609.36357] HyperZip: Efficient Data Compression through Personalized Diffusion LLMs with Hypernetworks (https://arxiv.org/abs/2609.36357)
- HyperZip, full text on arXiv (https://arxiv.org/html/2609.36357)
- [2608.11249] Diffuse to Compress: Leveraging Diffusion LMs for Lossless Compression (https://arxiv.org/abs/2608.11249)
- [2609.36938] Efficient Agentic LLM Serving over SSD-based Sparse KV Storage (https://arxiv.org/abs/2609.36938)
- [2609.36505] BRIDGE: Bilevel Retrieval-Credit-Aware Agentic Reinforcement Learning (https://arxiv.org/abs/2609.36505)
- [2609.37052] OmniRoute: Mapping Temporal Semantic Evidence to Audio-Visual Token Budgets for Efficient Omnimodal Large Language Models (https://arxiv.org/abs/2609.37052)
- [2609.37655] Exemplar2VQA: A Scalable Exemplar-Driven Visual Question Answering Generation Framework via Multi-Agent Coding (https://arxiv.org/abs/2609.37655)
- Inception Launches Mercury 2.5, the Next Tier of Intelligence for Diffusion LLMs (https://sg.finance.yahoo.com/news/inception-launches-mercury-2-5-163000864.html)
- Mercury 2.5: API Provider Benchmarking and Analysis, Artificial Analysis (https://artificialanalysis.ai/models/mercury-2-5/providers)
- Introducing System One Models and Jev, TypeSafe AI (https://typesafe.ai/blog/introducing-system-one-models-and-jev)
- browser-use/jev-ultrafast, GitHub (https://github.com/browser-use/jev-ultrafast)
- NandhaKishorM/laya, GitHub (https://github.com/NandhaKishorM/laya)
- mizorewww/laya-mlx, GitHub (https://github.com/mizorewww/laya-mlx)
- jaredpalmer/kev, GitHub (https://github.com/jaredpalmer/kev)
- tamaratran/fast-jev-compaction, GitHub (https://github.com/tamaratran/fast-jev-compaction)
- LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference, MLSys 2026 (https://mlsys.org/virtual/2026/oral/3646)
- Precise Prefix Cache Aware Routing, llm-d (https://llm-d.ai/docs/guides/precise-prefix-cache-aware)
- DDN Accelerates Inference and Lowers Cost Per Token While Expanding Multi-Tenant Training for AI Factories (https://www.ddn.com/press-releases/ddn-accelerates-inference-and-lowers-cost-per-token-while-expanding-multi-tenant-training-for-ai-factories/)
- Senior Software Engineering Manager, KV Cache Platform, DDN Careers (https://careers-ddn.icims.com/jobs/5967/senior-software-engineering-manager-%e2%80%93-kv-cache-platform/job)
- SOSP 2026 Accepted Papers (https://sigops.org/s/conferences/sosp/2026/accepted.html)
- PyTorch Conference 2026 (https://pytorch.org/event/pytorch-conference-2026/)
- PyTorch Foundation Welcomes vLLM as a Hosted Project (https://pytorch.org/blog/pytorch-foundation-welcomes-vllm/)
Disclaimer: This article is generated by GatiFlow Intelligence for informational purposes only. It does not constitute investment advice, recruitment recommendations, or legal guidance. All data is derived from public sources and AI analysis — verify independently before making decisions. Past trends do not guarantee future results.
Where this came from
Every Deep Dive starts from GatiFlow's own pipeline: 13 public developer sources, collected every six hours, with a confidence score and the evidence behind each signal. The same signals, filtered to the topics you follow, are a JSON API.
No credit card required.
Get the next one by email
One article every Saturday morning in your time zone. No account needed, and nothing else is sent to the address.
Double opt-in: you confirm by email first. What we store, and for how long, is in the privacy policy.
Tell me I am wrong
Corrections, the version of this you have lived through, or what you would like covered next. It reaches me directly and is never published. It is kept for two years so it can be read and answered; the privacy policy has the details.
0/2000