Conversation
Notices
-
Embed this notice
apparently giga llms are just mixtures of 9b models. so fine tuning your own 9b model to a specific task should be reasonably frontier grade, asymptotically. weird.
-
Embed this notice
@meeper yeah. mixture of experts is basically 9B models crammed together with a "router" that decides who is active.
which means 9B is about where the labs have converged on the model doesn't get more useful above this size.
-
Embed this notice
@icedquinn you're talking about MoE stuff right?