Loading...
Loading...
Loading...

EverMind AI Sets the New Standard of RAG with EverMemReRank

EverMind AI Sets the New Standard of RAG with EverMemReRank

Part of the EverMemModel modules, the ReRankModel achieves SOTA performance on 2wiki and Hotpotqa.

EverMind researchers

About 1 minutes to read

SOTA
2wiki
Hotpotqa
RAG
EverMind AI Sets the New Standard of RAG with EverMemReRank

We have integrated a generative ReRankModel into the traditional RAG retrieval framework. By this week, we find it on the right track of performance improvement on key benchmarks.

Performance on 2wiki benchmark

When compared with HippoRag2 under identical conditions using Llama3.3 as the LLM for QA, our method achieved a QA F1 score of 0.758 on the 2wikimultihopqa public benchmark, outperforming HippoRag2 by 4.8 percentage points and reaching SOTA level.

table 1

Performance on Hotpotqa benchmark

On the HotpotQA public leaderboard, our model achieves an F1 score of 0.7802, outperforming HippoRag2’s 0.755 by 2.5 percentage points and reaching SOTA level.

table 2

We have integrated a generative ReRankModel into the traditional RAG retrieval framework. By this week, we find it on the right track of performance improvement on key benchmarks.

Performance on 2wiki benchmark

When compared with HippoRag2 under identical conditions using Llama3.3 as the LLM for QA, our method achieved a QA F1 score of 0.758 on the 2wikimultihopqa public benchmark, outperforming HippoRag2 by 4.8 percentage points and reaching SOTA level.

table 1

Performance on Hotpotqa benchmark

On the HotpotQA public leaderboard, our model achieves an F1 score of 0.7802, outperforming HippoRag2’s 0.755 by 2.5 percentage points and reaching SOTA level.

table 2
Loading...
Loading...

You may also like these

Related

SkillCorpus paper title card

Do public SKILL.md files actually make agents better?

SkillCorpus,SKILL.md,agent skills,skill curation,skill retrieval,LLM agents,SkillsBench,GDPVal,agent harness

CRAFT paper title card

CRAFT: learning how to fuse video tokens, not just which to drop

CRAFT,video token compression,vision-language models,video VLM,KV cache,prefill cost,token merging,token pruning,temporal reasoning

HarnessBank paper title card

Self-evolving agents have a measurement problem

self-evolving agents,agent harness,HarnessBank,credit assignment,LLM agents,agent evaluation,harness optimization,significance testing

Skill Hub: a measured foundation for community-powered agents

skillhub,skill benchmark,SKILL.md,community skills,ai agent

EverMind AI Sets the New Standard of RAG with EverMemReRank

Part of the EverMemModel modules, the ReRankModel achieves SOTA performance on 2wiki and Hotpotqa.

EverMind researchers

About 1 minutes to read

SOTA
2wiki
Hotpotqa
RAG

EverMind

A straightforward solution to long-term coherence

Scan to join the community

Discord

Wechat

© 2026 EverMind Team.

EverMind

A straightforward solution to long-term coherence

Scan to join the community

Discord

Wechat

© 2026 EverMind Team.

EverMind

A straightforward solution to long-term coherence

Scan to join the community

Discord

Wechat

© 2026 EverMind Team.