seeniceism

Asynchronous. Inference. Primitives.

I started this essay to introduce what I was building, but then I had to stop and go back to write two more (here and here) before publishing this. I had to capture the true intent and all the pivots in thinking, building and the evolving hypotheses behind what eventually turned out to be a set of asynchronous inference primitives.

In case this essay reached you first, here is the summary of the prequel.

So I was a new dad wanting to get the most out of AI but interleaved with changing diapers and bottle feeding. Chat sometimes was too much work for little flow in return.

The shape and size of my time were changing, and so was the work I was submitting to the models.

I was hoarding screenshots and wanted a way to synthesize them with AI and seamlessly surface and repurpose results whenever I needed. But small models and my ancient hardware made it hard to implement such workflows. Later when the ecosystem turned agentic the frameworks and API wrappers were too bundled. Powerful, but too big for my needs especially for working locally.

Then I figured out that most of the bundling was there to be compatible with the OpenAI-style v1/chat/completions REST API. So everything mostly was chat by default and other asynchronous-looking solutions also mostly were running a session with the model underneath. I found no alternatives to these synchronous primitive APIs.

So I pivoted from changing interfaces and managing screenshots to filling what I thought was a gap with asynchronous inference primitives.

Now back to this essay.

My first attempt at turning my abstract ideas into code and solving some personal pain points locally was called pnpl. Push Now and Pop Later (naming is hard even outside of Computer Science!). And the good thing about wanting to run local models on older hardware without a GPU was that I got introduced to llama.cpp! The impact of that open source runtime and its community is immeasurable. llama.cpp also continues to be the only constant through the changing phases of this side quest.

So when I came back the second time with the same hardware and different intent, the small model ecosystem was not small anymore.

The combination of Hugging Face and the llama.cpp community was pushing small models to the frontier!

Me and my hardware were set up for success with access to an ever-increasing number of the latest and multimodal models and the ability to run them seamlessly even on my old laptop.

For the first time I was getting work done with the primitives I was building to run local models that suited my 2017 MacBook Pro with an Intel Chip.

As they say, nothing works like work. Once I started to see results I could use and the experience of using multiple multimodal models to run custom workflows , my center which drove all judgement and choices and assumptions to build, shifted.

And I believe what you put in the center determines what gets built around it.

It is then that I started to see the current ecosystem as a galaxy with the big synchronous chat at its core. Over time, there were smaller systems formed within it around their own centers, for example coding, agentic sessions, model runtimes and enterprise AI.

Each such smaller system gave rise to its own orbit of tools, software, hardware, metrics, frameworks and economics, all conforming to its local center while still living within the larger synchronous chat based galaxy. So finding chat and chat based frameworks everywhere was a feature, not a bug!

All this led me to again change my hypothesis.

Probably there are no gaps to fill, only centers to find and build around.

Once I stopped looking for gaps and looked for centers instead, it was easier to see why I could not find the primitives I wanted earlier, and why even building them kept getting complicated.

I was looking in the wrong orbit. Maybe even in the wrong galaxy!

So roughly after three summers for me, during this second stint with building primitives, I tasted disproportionate flow for the first time. All this momentum and increase in build velocity was because I had a new center, outside of the synchronous galaxy, it was

the work I wanted done asynchronously!

Yes, everything I was building for and solving toward now was for my ever evolving *work loads. This recentering reinforced that when work is at the center, the intelligence can stop solely influencing the ecosystem. It should continue to be powerful but served differently. Such that all work does not have to be in a session by default to access AI. The best generally available intelligence for the hardware and the work should just be provided like a utility.

I did not want the lifetime of my work and its results to be tied to the runtime of a model.

Once intelligence became utilitarian, I could bring focus back to the center, my work. Ideas emerged about how I wanted to submit work as it arrives or in bulk without getting blocked. How I wanted the freedom to access results and repurpose them. Interesting questions emerged too,

But again I have been cautious not to end up bundling a queue service, a database, another server, timers and retries just to make the workflow seem asynchronous. I have been relying on the ideal stack already on my machine and everyone else's!

Files. Directories. Processes.

Filesystems have worked robustly for decades, letting the complexity, if any, build on top of them. They have stayed simple, swappable, small and incredibly useful.

So letting go of older orbits, recognizing the right center and looking closely at what always worked led me to find

nrvna.

A set of asynchronous inference primitives to make work and results durable.

nrvnad - run a model against a workspace

wrk - submit work to a workspace

flw - retrieve results from a workspace

I am cleaning up the repo and rewriting some docs, so you and your agents can build in this new galaxy centered around the work you want done asynchronously.