Hacker Newsnew | past | comments | ask | show | jobs | submit | CRG's commentslogin

It's also smaller than GPT-2 (1.2B vs 1.6B) and trained with a lot less language data (6% of the training mix).


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: