hosted personal AI assistants with persistent memory: comparing Jenova, Lindy, Vellum, and similar platforms
Persistent memory separates AI tools people abandon from ones they actually keep using.

Memory is the real dividing line among AI assistants right now, ahead of price, the model card, or how slick the chat window looks. Does the thing remember you tomorrow, or does every session start from zero? I've spent the last few months poking at both ends of that question, and it's the one thing almost nobody checks before signing up for another tool.
Live with a stateless assistant for a few weeks and the gap gets obvious fast. It gives you a sharp answer in the moment, then forgets your job, your preferences, your half-finished project the second you close the tab. A persistent assistant carries context forward. Tools like Vellum, a memory-powered AI assistant that learns your preferences over time, are built around exactly this distinction. You stop re-explaining yourself every morning, and that sounds small until you've done it forty times and started to resent it a little.
One thing worth clearing up: a bigger context window delays the reset but doesn't prevent it. Whatever the model picked up inside that window disappears once the session ends, no matter how roomy it felt while it lasted. I think about this in three tiers. Short-term context lives inside one open chat and dies when you close it. Session-based memory survives a workflow, maybe a few tool calls strung together, but resets the next time you log back in. Persistent long-term memory survives across weeks and pulls up what's relevant, and that's the tier that actually changes how you work day to day.
I went looking at the numbers because I wanted to know if this was more than a semantic argument people have on Twitter. The AI agent memory segment hit $6.27 billion in 2026, projected to reach $28.45 billion by 2030. Money doesn't move that fast unless stateless agents were failing people in a way that cost something real: hours, redone work, the friction of explaining yourself to a machine for the third time this week.
So what actually separates a real memory system from a slide in a pitch deck? Storage type, for one. Conversation history, stated preferences, project context, recurring patterns: these need to live in separate buckets and get pulled back appropriately, rather than dumped into one undifferentiated pile. Control matters more, honestly. Can you see what the assistant has stored about you, edit it, delete it? Or are you just trusting a black box and hoping for the best? And cross-channel consistency gets underrated constantly. Memory that follows you from desktop to phone to Slack is doing its job. Memory that resets the second you switch devices is only performing the idea of memory.
The two platforms below made genuinely different bets on where memory should live and what it should hold onto. Those bets, more than pricing pages or feature checklists, decide whether the tool still earns a spot on your phone six weeks from now.
The broader market these platforms are competing in
The personal AI assistant market sat at $4.84 billion in 2026, projected to hit $19.63 billion by 2030. That's a 41.9% compound annual growth rate, something like five times growth in five years. Numbers like that pull in serious builders. They also pull in a lot of noise, and from the outside it's not always easy to tell which is which.
Adoption stopped being theoretical a while back. McKinsey's 2025 State of AI survey puts regular AI use at 88% of organizations in at least one business function, up from 78% the year before. Gartner projects 40% of enterprise applications will include task-specific AI agents by the end of 2026, up from under 5% in 2025. That's the shift from experiment to infrastructure, happening while we're all still arguing about whether any of this is overhyped.
Even the biggest player in the space is losing ground, and the product didn't get worse. ChatGPT held around 87% of global generative AI traffic at the start of 2025. By year's end, that number had dropped to roughly 68%. People didn't leave because ChatGPT declined. They left because they figured out different models are better at different jobs, and once you know that, it's hard to unlearn it.
That fragmentation is the world this whole comparison lives in. A crowded, fast-growing market means platforms earn loyalty through depth, not polish, and persistent memory is what turns a tool people try once into a tool people actually keep.
Jenova: multi-model access unified under one persistent memory layer
Jenova comes from Azeroth Inc., launched in August 2024, and it's already running in 60 countries. That's quick traction for something barely into its second year.
The idea is simple to state: instead of picking one frontier model and living with whatever that model happens to be weak at, Jenova routes each query to whichever model fits it best. Complex reasoning goes to Claude Opus 4.5, creative writing goes to GPT-5.2, fast and cheap tasks run on Gemini 3 Flash. If you'd rather pick yourself, manual model selection is right there too.
None of that matters much if memory falls apart across the switching, though. Unlimited persistent memory ships on the free tier, no premium upgrade required. That tells you something about how central memory is to the pitch, rather than a footnote feature bolted on later.
What actually lives in that memory layer? Preferences, recurring patterns, project context, carried across every agent and every model call. Talk to the Academic Research Assistant on Monday, the Business Co-Pilot on Tuesday, a custom agent you built yourself on Wednesday, and the memory follows you, rather than staying tied to whichever agent happens to be running that day.
On integrations, Jenova connects to over 200 apps through the Model Context Protocol (MCP), including Gmail, Google Calendar, Google Drive, Notion, Dropbox, Reddit, and YouTube. Since it's MCP-compliant, that list grows to any compliant server going forward, so the integration surface isn't a fixed asset that shrinks in relevance over time.
Paid subscribers can build custom agents too: personalized instructions, private knowledge bases, a chosen model, specific tool configurations, all running on the same infrastructure as the first-party agents. The first-party lineup holds its own on that front, with an Academic Research Assistant that has Google Scholar built in, a Business Co-Pilot, a Technical Stock Analyst, a Travel Planning Advisor, and a Creative Fiction Writer.
On privacy, Jenova states user data isn't used for model training or marketing. Pricing starts at $20 a month for paid plans, and the free tier already includes the memory layer.
Here's what I keep coming back to, though: model routing is only as good as the platform's judgment about which model fits a given query. Memory is only as good as its ability to carry context across those switches without losing fidelity. That gap between the pitch and the day-to-day experience is the kind of thing you test yourself. You won't settle it by reading a comparison article, including this one.
Lindy: ambient execution with per-agent memory and a notable scope boundary
Lindy repositioned itself as a personal AI assistant with SMS and iMessage built in. That's a real pivot, showing up inside the channels you're already using to talk to actual people, rather than asking you to open a new app.
Reachable through email, iMessage, Slack, and a browser, it feels ambient rather than stuck off in its own destination somewhere. It has drawn consistent praise for fitting naturally into daily communication workflows rather than demanding you learn a new interface. A tool that fits inside habits you already have gets used far more than one that asks you to build new habits from scratch, so that reputation is earned.
The memory architecture holds up well within its lane. Per-agent memory persists across runs: the email agent remembers the customer, the last thread, your preferred tone. That memory is editable and resettable from the dashboard, and every step shows up in a per-run log, so you're not left guessing what happened or why.
There's a documented boundary here, though, and it's worth saying plainly instead of burying it. The core knowledge layer is session-scoped, meaning each new chat starts fresh by default. There's no persistent organizational memory across conversations in the base setup, and that limitation shows up again and again in third-party reviews. In practice, Lindy remembers your email habits inside the email agent just fine, but it doesn't build one unified picture of you across everything else you're doing elsewhere. The memory runs deep within its lane and stops at that lane's edge.
On integrations, Lindy claims over 5,000 as of early 2026, covering most of the SaaS tools small and mid-market teams already lean on daily.
Pricing got restructured in early 2026, and the free plan disappeared entirely. Plus runs $49.99 a month for standard usage, up to two inboxes. Pro is $99.99 a month for three times the usage, three inboxes, computer use, and model selection. Max runs $199.99 a month for seven times the usage and five inboxes.
There's reportedly an enterprise onboarding fee around $1,500 that doesn't show up on the main pricing page, and actual monthly spend tends to climb once voice calling, multiple phone numbers, or high-frequency workflows enter the picture. Dedicated support is Enterprise-only; Plus through Max get help center and community support instead, and there's no annual discount currently.
Lindy has raised over $50 million as of mid-2026, which makes it a funded company defending a roadmap rather than a side project testing an idea over a weekend. The best fit here is a single professional who needs one job done well, calendar triage or email drafting say, with fairly predictable volume and no real need for the assistant to connect dots across unrelated corners of their work.
How memory architecture differs between platforms built for individuals versus those built for workflows
Two design philosophies keep surfacing across this comparison. One is workspace-wide memory: a single picture of the user that persists across every agent, every model, every tool, so the assistant learns you as a whole person rather than a task. The other is agent-scoped memory: each specialized agent holds its own context, built for depth in one job, without necessarily sharing a layer underneath it.
Agent-scoped memory tends to win when the task is bounded and repeats constantly: same inbox, same customer list, day after day. Workspace-wide memory tends to win when the work is varied, when context bleeds across domains, and what you actually want is an assistant that connects things over time instead of treating each task like its own island. Neither one is the obviously correct default.
Here's the engineering wrinkle I keep circling back to. If a platform routes queries to different models depending on the task, the memory layer has to hold your context steady across those switches. That's a genuinely harder problem than memory built for one fixed model, since the system has to translate context consistently no matter which model happens to be reasoning underneath it at that moment.
User control turns out to be a real differentiator even between platforms with strong memory on paper. Can you actually see what's stored about you, and fix something wrong without starting the relationship over from scratch? Or is the memory a background process you're just supposed to trust blindly?
Self-hosting is its own trade-off entirely. Going open-source and running your own infrastructure kills platform dependency and keeps data on hardware you control, but it demands overhead most individual users won't take on willingly. That math works better for teams with real governance requirements than for someone just trying to manage a calendar and an inbox.
Cross-channel consistency might be the most overlooked piece of all this. A memory that forgets you the moment you switch from web to phone isn't offering much, even if it looks sophisticated sitting on your desktop screen.
What to weigh when choosing between these platforms for your actual use
Start with the memory question before you even open the feature list. Do you need the assistant to know you across a lot of different kinds of work, or to know one domain cold? Is a unified picture across your whole workspace worth more than deep memory inside a single task? That answer decides most of what follows.
Channel access matters more than it usually gets credit for on a spec sheet. If the assistant lives in a separate app you have to remember to open, you'll use it less; that's just how habits work. Lindy's SMS and iMessage integration solves that directly. If your work is already scattered across a pile of different tools, a platform with broad MCP-based integration, Jenova's 200-plus apps being the case in point, cuts down on the work of wiring everything together yourself.
Check pricing against your actual usage, not the number on the landing page. Lindy dropping its free tier, plus the gap between the sticker price and what you'll actually pay once voice and high-frequency workflows get added in, makes it harder to try before committing real money. Jenova's free tier includes the full memory layer, which lowers the bar for figuring out whether its memory model fits your work before you spend a dollar.
There's a real trade-off between flexibility and simplicity too. Multi-model routing, like Jenova's, is strong if you trust the system's judgment, and it can feel unpredictable if that routing logic stays opaque to you. Single-model execution, like Lindy on its Plus tier, is more consistent, especially for bounded, repeatable tasks where you don't need the system making judgment calls about which model fits best.
A few privacy questions belong on the table for any platform in this category, not just the two here. Is your data used for training? Can you export or wipe your memory store whenever you want, no questions asked? Where does the data actually live, and does self-hosting change that math if governance is a real concern for your team?
Test the memory before you test anything else. The chat window sitting on top of it won't tell you much; the memory underneath it will tell you everything.


