Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Does anyone know if the attention mechanism truly does consider all of the context or does the model use some of the recent tricks for reducing the context size before running the attention process?


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: