Johannes Uusikuu's picture

Johannes Uusikuu

Johneeee

AI & ML interests

Small omlx models without vision & mtp for m1 m2 or low memory macs. Try the oQ63e quants they have 5 bit base but bump a lot of the model to 8 and 6 bit. oQ63e seems to work very well. My experimentation shows that giving a big bit budget bumps up 24 % of the tensors to 8 bit and around 15 percent to 6 bit.

Recent Activity

Organizations

None yet