Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Is it a fully custom inference engine or are you reusing parts of an existing stack? (llama.CPP, vLLM, etc.)


Our skippy library is a patch queue on top of llama that allows us to access internal information, such as activations, and filter tensors on model load.


This really should be in the blogpost. It’s both useful info and basic courtesy to be explicit about which underlying inferencing engine you are using


We didnt post it, we use a library (iroh) who featured us - so we are here answering any Q’s instead :)




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: