You got me interested, a rough table of model size to instance type to spot $/hr would help a lot. The 0.5B example is CPU only, so it doesn't say much about what a 7B or 70B actually costs.
You're right. A CPU only model is definitely is not a real prod workload where you'd need GPU machines.
Here is a rough table of model size to instance type to spot $/hr:
For up to 8B, you can use a c6i.xlearge which costs as spot about $0.4/hr.
For up to 70B, you can use a g5.12xlearge that costs you about $2/hr.
Besides the GPU machine be aware that you need a controller instance to route the requests and scale. That has a fixed cost of up to $0.08/hr. When not in use at all, just type veloxml down --all and it tears down the controller too for true $0/hr.
We're currently running larger LLM models benchmarks to add to the README this week.
Great question, and I just updated the README to clarify this! It's a mix. If you're using connected agents (Antigravity, Claude Code, Cursor) via MCP, they have the engrim_add tool and will automatically log major architectural decisions as they make them. But you also have the manual engrim add CLI command for when you have an 'aha!' moment and want to drop a constraint into the project's brain yourself.
Nice concept, the offline provisioning is the part that stands out to me vs Tailscale. How does revocation work in that model? Didn't see it in the docs, if a laptop with a .peer profile gets stolen, what then?
Just at the right time, will be trying this out.
I see it as an addition to local LLM, where on a job cycle it can learn from the LLMs contexts and free it up and help it in the more complex work.
First some kind of before and after benchmark, so you can actually see if defragging made a difference.
Second is a recommendation for defragging based on things like the fragmentation, free space, and whether it’s an HDD or SSD.
Especially the second one feels useful since defragging doesn’t necessarily mean improvement to the system as you said yourself.
It does work partially. TokenSift can analyze the textual parts of prompts sent to multimodal models, but doesn’t calculate or optimize image/audio/video token usage. Non-text payload calculation is something I’d like to add.
Interesting concept, I really like the idea of monitoring the domain.
I tried 1 domain I am looking at: bein.com, resulted in a thousand domains and not this one. Will give it a better look though.