“The next generation of edge AI will require more than increasingly powerful models. It will require an entire systems stack designed around the realities of the physical world: constrained compute, intermittent connectivity, real-time response, privacy, power efficiency—and memory.” – Tara Khani, CEO and Founder
As AI evolves from models that simply respond to prompts into agents that perceive, reason, remember, and act, a new infrastructure challenge is becoming increasingly important:
How do we give AI reliable memory?
The EDGE AI FOUNDATION is excited to welcome Moorcheh.ai by Edge AI Innovations to our growing global ecosystem of organizations building the technologies that will bring intelligent systems closer to where data is created and decisions are made.
Moorcheh is tackling one of the less visible—but increasingly critical—layers of the AI stack: retrieval and memory for AI agents.
The Next AI Bottleneck May Not Be the Model
In a recent conversation on The Makers, Moorcheh CEO Tara Khani and CTO Majid Fekri explored an important shift taking place across agentic AI.
As models become increasingly capable, the reliability of an AI system depends on much more than the model itself. Agents need to remember previous interactions, retrieve the right information at the right moment, distinguish new information from outdated information, resolve contradictions, and maintain useful context over long-running workflows.
That makes memory and retrieval a fundamental systems problem.
Many of today’s RAG and semantic-search architectures were originally designed around approximate vector retrieval. That works well for many search applications, but autonomous agents introduce a different requirement: an agent may repeatedly retrieve information, make decisions based on it, add new information to memory, and use those decisions as context for future actions.
In that environment, inconsistent retrieval can become inconsistent behavior.
Moorcheh is approaching this problem from first principles.
Rethinking Vector Search Through Information Theory
Instead of relying on the conventional combination of high-precision floating-point vectors, HNSW indexes, and geometric similarity measures, Moorcheh has developed an information-theoretic retrieval architecture.
At the core is Maximally Informative Binarization (MIB), which transforms embeddings into compact binary representations designed to preserve semantically useful information.
Combined with Moorcheh’s Efficient Distance Metric and Information-Theoretic Scoring, this architecture enables the system to perform deterministic retrieval using highly efficient bitwise operations.
The implications extend beyond search performance.
Smaller representations mean lower memory requirements. Efficient retrieval means less compute. And eliminating dependence on large in-memory vector indexes creates new possibilities for deploying semantic search and AI memory in environments where conventional cloud-oriented architectures may be impractical.
That is where Moorcheh becomes particularly relevant to the edge AI community.
Bringing Retrieval and Memory to the Edge
Edge AI is fundamentally an exercise in resource allocation.
Models, embeddings, retrieval, sensor processing, application logic, and increasingly agentic workflows must share limited compute, memory, storage, and power.
If retrieval consumes the resources needed by the model, truly local agentic AI becomes difficult.
Moorcheh has demonstrated its retrieval architecture running directly on constrained hardware, including a local RAG implementation on Arduino UNO Q. In its published testing, Moorcheh reports approximately 17 ms median semantic search across 10,000 vectors while using roughly 20 MiB of server RAM.
This creates an intriguing possibility:
What if an edge device could not only run AI locally—but remember locally as well?
- A robot could retain operational knowledge without continuously querying the cloud.
- An industrial system could retrieve maintenance and diagnostic information inside the facility.
- A retail kiosk could understand its local catalog while keeping interactions private.
- A field device could maintain contextual knowledge even when connectivity disappears.
- And personal AI systems could retain useful long-term context while keeping sensitive data within the user’s own environment.
The edge, in other words, can become more than the place where inference happens. It can become a place where AI maintains state, context, and memory.
From Search to Agentic Memory
Moorcheh is extending this work through Memanto, its memory layer for long-running AI agents.
Memanto organizes agent memory into typed semantic categories while incorporating capabilities such as temporal versioning, source attribution, and conflict resolution.
The objective is to move beyond simply storing conversation history toward infrastructure that helps agents manage what they know over time.
That distinction becomes increasingly important as the industry moves toward persistent agents, multi-agent systems, robotics, physical AI, and other applications where AI systems operate continuously rather than within isolated prompt-response sessions.
Joining the foundation
As we bring together the hardware, software, research, developer, and application ecosystems required to move AI from centralized infrastructure into the real world. Moorcheh adds an important piece to that stack: efficient retrieval and persistent memory.
Its work sits at the intersection of several priorities across our community:
- Efficient AI — reducing the compute and memory overhead associated with semantic retrieval.
- On-device intelligence — enabling RAG and memory architectures to operate directly on edge hardware.
- Privacy and data sovereignty — allowing knowledge and memory to remain inside a device, facility, private cloud, or sovereign environment.
- Agentic and Physical AI — giving increasingly autonomous systems persistent access to relevant context.
- Developer accessibility — making these capabilities available across cloud, on-premises, and edge deployment models.
Moorcheh has already been showing up quite actively in the EDGE AI FOUNDATION programming. The most relevant recent participation:
- Sept. 24, 2026 — EDGE AI Workshops: “From Language Models to Sensors.” Moorcheh CTO/co-founder Dr. Majid Fekri co-hosted “Building a Privacy First Voice Kiosk: Local RAG & LLMs on Arduino UNO Q” with Qualcomm/Arduino. The workshop demonstrated a fully local assistant combining Moorcheh Edge’s 1-bit embedding compression and semantic search with Llama 3.2 1B, Whisper and Piper—all without cloud APIs. Arduino also offered UNO Q boards and attendees received complimentary Moorcheh Edge access.
- Recent Generative Edge AI programming — “Total Recall at the Edge.” Tara Khani of Moorcheh.ai presented “Total Recall at the Edge: A Deterministic Universal Memory Layer for Agentic AI,” connecting Moorcheh’s work directly with the Foundation’s Generative Edge AI / agentic AI focus.
We are excited to welcome Moorcheh.ai by Edge AI Innovations to the EDGE AI FOUNDATION and look forward to working together to explore what becomes possible when AI can not only reason at the edge, but remember at the edge.
Welcome to the EDGE AI FOUNDATION, Moorcheh.ai!
