Follow

Jay Alammar's overview of BERT-based models shows us some not so obvious conclusions - like the fact that the best way of getting a contextual word embedding isn't taking the output layer of the model.

jalammar.github.io/illustrated

Sign in to participate in the conversation
CleverLibre Social

CleverLibre Social is an inclusive social instance for open discussion, learning, and community.
All cultures welcome.
Hate speech and harassment strictly forbidden.