Google 发布 EmbeddingGemma 2,称多模态嵌入基准超越两倍规模竞品
Google 发布开源多模态嵌入模型 EmbeddingGemma 2,可将文本、图像、视频、音频和代码转为数值向量。该模型 7.4 亿参数,在 Massive Text Embedding Benchmark(Code)上得分 78.68,比前代 68.76 提升近 10 分,Google 称其表现超过规模两倍的竞品。
Google 发布开源多模态嵌入模型 EmbeddingGemma 2,可将文本、图像、视频、音频和代码转为数值向量。该模型 7.4 亿参数,在 Massive Text Embedding Benchmark(Code)上得分 78.68,比前代 68.76 提升近 10 分,Google 称其表现超过规模两倍的竞品。
H Company 团队发布 NeoMME,含 260M 和 800M 两个尺寸的多语言多模态编码器,用单个双向 Transformer 从零处理文本 token 和 32×32 图像 patch,不使用预训练视觉塔或因果语言模型,训练采用掩码离散扩散目标,每个模型处理约 5240 亿 packed token。