Wire
LightOn trains 307M retrievers on 2.8B pairs
LightOn released mDenseOn and mLateOn, two open 307-million-parameter multilingual retrievers trained on a 2.8-billion-pair corpus, alongside the models, datasets, and training code in its technical release. On MIRACL languages absent from retrieval training, mLateOn averaged 67.59 versus mDenseOn’s 57.42, a 10.17-point spread; for teams following the expanding open-weight model ecosystem, the result argues for testing late interaction—not just adding parameters—when retrieval must cross languages and scripts.