FinPerMA Benchmarks LLM Agent Personalized Memory in Finance

Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang· August 6, 2026 View original

Key takeaways

  • LLM agents struggle to maintain and update personalized user models over long horizons, especially in finance.
  • FinPerMA is a new benchmark for evaluating event-driven preference adaptation in financial LLM agents.
  • Current frontier LLMs show low accuracy in personalization, particularly after significant financial events.
  • Simple retrieval can sometimes outperform purpose-built memory systems for personalization.

Who benefits

BFSIFinTechWealth ManagementInsurance

Summary

Researchers introduce FinPerMA, a new benchmark evaluating Large Language Model (LLM) agents' ability to maintain and update personalized user models over time, especially in high-stakes financial advising scenarios. It tests event-driven preference adaptation against frozen longitudinal investor trajectories, revealing that current frontier LLMs struggle significantly with personalization, particularly after material financial events.

As Large Language Model (LLM) agents are increasingly deployed as personalized assistants, particularly in critical sectors like financial advising, their capacity to maintain and adapt individualized user models over extended periods is a key concern. Existing benchmarks for personalized memory often focus on factual recall or rely on loosely constrained, model-generated scenarios, leaving the crucial aspect of event-driven preference adaptation largely unexplored. To address this, FinPerMA has been developed as an event-grounded benchmark. It assesses personalized memory by testing LLM agents against pre-defined, longitudinal investor trajectories, incorporating deterministic, theory-informed impact rules and controlled LLM narration. A "Post-Shock" checkpoint specifically evaluates whether an agent successfully integrates significant financial events into its persistent user model. Evaluations across 2,994 questions from 276 personas, involving seven frontier LLMs and various memory configurations, show that current models are far from proficient. No full-context configuration achieved more than approximately 47% overall accuracy, or 39% on multiple-choice questions. Analysis indicates that summary-based memory often retains facts but loses critical preference signals needed for true personalization, sometimes making simple retrieval more effective, especially after market shocks.

Why it matters

For financial professionals and AI developers, understanding LLM agents' limitations in personalized memory is vital for building trustworthy and effective AI advisors that can genuinely adapt to individual client needs and market changes.

How to implement this in your domain

  1. 1Assess current LLM agent capabilities for personalized financial advice against the FinPerMA benchmark's criteria.
  2. 2Investigate advanced memory architectures beyond simple summarization for LLM agents in financial applications.
  3. 3Develop strategies to improve LLM agents' ability to integrate material events and adapt user preferences.
  4. 4Pilot new LLM agent designs with a focus on robust, long-term personalized memory for client interactions.

Original post by Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang

"arXiv:2608.04095v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long…"

View on X

Originally posted by Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses