GAI before AGI?
I am drafting a post to introduce something I have built. I wrote the title
Asynchronous. Inference. Primitives.
and I knew, before a landing page there has to be a "Why and How did I even land here?" page.
Let me take you back to the past where I encountered each of those words. Each of these time periods has a role in forming my thoughts around what to build and more precisely what not to build.
Back to the Future
It is mid 2024. AI is democratized, thanks to both Chat & GPT in the ChatGPT app. GPT is the brain behind serving useful AI to the consumers in real time. And Chat is of course chat. The default surface, easy, convenient and familiar. So AI chatbots are proliferating exponentially. There is almost no obvious reason to look beyond chat interfaces.
Except. I have two.
The first one is 2. The other has just turned a month old.
But I am not looking for alternatives to Chat interfaces simply because I am a new dad or I have less time or I do not like chatting. The real issue is more layered.
I am making the most out of my paternity leave. Some days a 1 hour window becomes available at 4am in the morning. Then there are days when a supposedly 3 hour downtime is done in 45 minutes. Other days I barely touch a device except for a few minutes in between.
So the shape and the size of my time is changing. It is unstable, uncertain and uneven.
All this stop and go got me thinking beyond the default way of using AI.
There surely seems to be room for a real life AI alongside the real time AI.
How can I always come back to a useful artifact? How can I hit flow state even with frequent stops? Without digging through tons of siloed chat histories to fast track myself back into a flow state. Can we have something not session based. One that adapts to the asymmetries of life. Some alternatives, more options so we are not always defaulting to synchronous conversational interface for everything AI.
Maybe something asynchronous?
Back to the Future Part II
It is somewhere near the end of 2024. Context window sizes are steadily increasing. I am relying on persistent chats for multiple Projects spread across my devices running a variety of models.
And with the multimodal models becoming mainstream, images mostly screenshots started dominating my prompts. Also models are turning into a commodity given their aggressive release cycles and ever increasing capabilities. All the options across open, free, local, thinking and reasoning variants, and my ever increasing number of screenshots, I am back to looking beyond the status quo again.
With more screenshots and more capable models two things happened
- I was running smaller sprints with a set of screenshots to achieve a specific outcome. And sometimes these had to be repeated. So this workflow based inference was unnecessarily hard inside chats.
- And then imagine the retrieval, repurposing and remixing. Some days I have an interview call tomorrow and I want a bunch of screenshots turned into flashcards to help me prepare. Other times there are two sets of screenshots which reinforce a concept I am learning.
So I started to believe such tasks are easier if I could somehow decouple prompts and inference.
Separate submission from execution? A solution where I had finer control over submission of prompts in a batch and probably retrieve them as one solid artifact? Push some jobs and pull results whenever time becomes available? Without scrolling or getting in and out of chat windows? Where artifacts, results and inference are easy to find, repurpose and retain?
Unlike last time when I chose changing diapers over interfaces, this time I decided to do something about these opinions I had been hoarding. I am building a small utility with the tech stack I know and understand.
I want it to be the simplest MVP decoupling inference from the chat window.
Back to the Future Part III
This is right around Summer 2025. Let me first acknowledge the elephant, I mean the Agent in the room. Chances are you are one or you might be using one to summarize this post and all while running a dozen others on your phones and terminals and browsers. Agents cracked writing code first and have not stopped.
And if you are wondering what happened to that little utility I was building around Christmas 2024, it did not fly very far. But building it got me deeper into small models and non conversational interfaces for local AI.
The models were still a problem though. There were not many reliable, performant small models that ran well on my 2017 MacBook Pro with Intel silicon. But now that seems to be changing. There has been tremendous improvement in the performance and availability of open, free, small models doing a lot of exciting things.
So you would think with ease of building agents and all these new small models one could just sit in a coffee shop and literally prompt oneself out of this corner. As I am onboarding myself with all these building blocks, primarily the libraries, tools, frameworks and SDKs, I am realizing that they are too big to be building blocks.
Because I have a specific problem I want solved involving screenshots and my laptop and small models, I am hoping to find something smaller, modular with the ability to grow with my solution.
Maybe a set of primitives?
Small enough to understand. Small enough to swap out. Small enough to keep only the parts I need.
The Future
Ok so we are in the Now(ish!), somewhere from the beginning of Summer 2025 till about whenever this post gets published.
Here is a hypothesis given I started local AI fairly early, on what is now ancient hardware. This was when local AI was not even close to being considered a default daily driver. Small models were not as performant.
So I did not roll out of my cloud accounts into my laptop just to give small models and Local AI a shot. Rather I was waiting there first with a well defined problem statement. Waiting for all the help I could get with the growing number of tools and libraries and frameworks developing within the Local AI infra ecosystem.
This order has proven to be important. Because it made it easy for me to see the bloat & bundling early on rather than in retrospect.
Most Local AI infrastructure seemed like a hand me down from the Cloud.
A lot of the dependencies, packages and solutions were built around chatting and coding. More context so you can chat more. More memory and harness engineering to reduce hallucination.
Many of the building blocks I encountered still assumed a server and a session underneath. And because there are not many alternatives to these server shaped AI interfaces, every app and every problem and corresponding solution inherits some of this baggage. With little freedom to opt out of defaults like GUI or context, caches, and memory related dependencies and packages.
I am aware this is all part of the price we pay to make persistent chats and long running tasks possible. Then the models can leverage all the context and provide a sense of continuity. Chatting with a model has never been my problem but that being the default for everything AI everywhere does become a problem.
Now for Local AI this lack of options becomes obviously inefficient at times. These chat heavy native apps, chat centric API wrappers and bloated building blocks are not setting me and my hardware up for success. Here I want alternatives to chat not just as a nice to have, but more so to leverage the strengths of Local AI.
Maybe I want it to work directly with my filesystem. Maybe I want the work and results to be durable and repeatable.
If I have a model locally downloaded and once I am sure of the prompts and the inference results, running it behind localhost cannot be the default.
I got to have a simple, non blocking way to get work done.
The Pivot
So I saw the bigger picture through this journey. And it helped me decide to build small.
Asynchronous. Inference. Primitives.
To get models to do your work without a chat window. And once you can do it reliably, repeatedly and durably, the intelligence you have locally becomes GA or Generally Available. In short GAI.
So maybe the better question is
GAI alongside AGI?