Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

There are some mad lads making different sizes of llama by "grafting" attention heads from one model onto another and finetuning a bit to stablize the transplant. For instance:

https://huggingface.co/models?sort=modified&search=20B

Its very experimental, but apparently the 20B models are actually improving on 13B.



Any place people doing such grafting are congregating?

I've often pondered if taking some random chunk of weights from the middle of a trained model, and dumping it into some totally different model might perform better than random initialization when the scale gets big enough.


I dunno. Probably the Kobold, Pygmalion, or AI Collective Discord, if I were to guess.

The first effort I am aware of is here: https://huggingface.co/chargoddard/llama2-22b

Being in the Discord age, lots of the discussion about cool llm stuff is fragmented and buried. I am in these discords, and I only know because I ran into it on HuggingFace.


Just the language being used here is amazing.


From the original effort linked above:

> This is Llama 2 13b with some additional attention heads from original-flavor Llama 33b frankensteined on.

And the lingo gets weirder, with an experimental merging method, for instsnce, being called "SLERP"


Is it something other than the standard meaning of "slerp" as spherical linear interpolation? That's not exactly new, it's been around since the mid-1980s.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: