does this work similar to airllm? i am wondering how it would handle something like quantizing kimi k3 on a budget of 8 gbs, or is that something you are not attempting to solve yet?
jaylane 2 hours ago [-]
tried it out but based on the model sizing result i got i got an insufficient memory error when the server started running