Product record
Release 3.6
Constellation, iCloud chat sync, Projects, a rebuilt chat drawer, and full-resolution image attachments
Noema 3.6 introduces Constellation, which lets you use your Mac's models from your iPhone, iPad, or Vision Pro with nothing to set up, over whichever connection is the most private one available at the time. Chats can now sync across your Apple devices through your private iCloud account, Projects group related conversations under shared instructions and sources, the chat drawer has been rebuilt, and images you attach finally reach the model at full resolution.
What changed
- 01
Constellation: use your Mac's models from your iPhone, iPad, or Vision Pro. Any Mac on your iCloud account with Remote Access on is found on its own - no address or token to enter - and tapping Use opens chat right away while the Mac loads in the background.
Highlight - 02
Noema picks the most private way to reach your Mac for each message, from your local network through to an encrypted relay, and shows which one it is using in the chat header. Off-grid Mode now holds across every kind of connection, not just ordinary web traffic.
- 03
Your Mac can stand in for the cloud: its models appear in Autopilot's stronger-model picker, and a new fallback model means a hard question tries another model before your on-device one answers.
- 04
Chats sync across your Apple devices if you want them to - one switch, off by default, in a new Sync & Devices setting. Conversations, bookmarks, and per-chat settings travel through your private iCloud account, and a reply that finishes on your Mac appears on your phone without opening the app.
Highlight - 05
Projects group related chats under shared instructions and sources, so every conversation in a project starts from the same context. Import PDFs, EPUBs, and text files straight into a project, or attach datasets you already have.
Highlight - 06
A rebuilt chat drawer brings search, projects, bookmarks, favourites, and date groups into one list with a preview of the last reply on every row, and chats stay noticeably faster during long generations.
- 07
Images you attach now reach the model at full resolution instead of being shrunk on the way in, so screenshots, receipts, and dense charts are legible to it for the first time.
Highlight - 08
Paged Overfit models run with the context you choose instead of a hidden cap, and Laguna S 2.1 now works as a paged model.
- 09
Models are more honest about what they can do: a vision model that cannot actually read images no longer claims it can, and downloaded models use the chat template they shipped with.
- 10
On Mac, the Relay page is replaced by a simpler Developer page, and Remote Access has moved into Settings, under Sync & Devices.
- 11
Interface fixes, bug fixes, and polish throughout.
Earlier releases
Open any record for its complete change set.
Release 3.5Noema Overfit, PDF attachment parity, longer conversations, and steadier model state
Noema 3.5 introduced the experimental Noema Overfit runtime for running compatible models far beyond normal device-memory limits by paging experts from local storage. PDF attachments began behaving consistently across iPhone, iPad, and Mac, while context compaction kept long conversations responsive. Model unload state, key/value cache reliability, and interface polish also received focused fixes.
Highlights
- Noema Overfit (experimental): run models far larger than your device's memory would normally allow — even 100B-class models on a laptop — by streaming their experts from storage on demand. This is early and evolving, so expect rough edges.
- PDF attachment parity: attach and read PDFs the same way across iPhone, iPad, and Mac, with consistent inline indexing everywhere.
- Longer conversations that don't stall: new context compaction keeps chats going smoothly as they grow, so Noema stays responsive deep into a thread.
- Fixed a model that could still appear loaded in Stored after it had actually been unloaded — for example when the app closed or the model was ejected unintentionally.
- Fixed the key/value cache occasionally disappearing during a session.
- Interface fixes, bug fixes, and polish throughout.
Release 3.4Document attachments on iPhone and iPad, a redesigned chat status drawer, steadier iCloud sync, tidier model settings, and sharper rendering
Noema 3.4 lets you attach PDFs, EPUBs, and text files to chat from the Files app on iPhone and iPad, with inline indexing so answers can use them right away. A redesigned chat status drawer makes capabilities, reasoning, Autopilot, MCP, and context usage clearer, while iCloud sync, model settings, math rendering, web citations, and model unload state all receive focused improvements. Interface polish, bug fixes, and localization round out the release.
Highlights
- Attach documents to chat on iPhone and iPad: add PDFs, EPUBs, and text files from the Files app and watch them index inline, so answers can draw on them right away.
- Redesigned chat status drawer: compact capability controls, reasoning and Autopilot status, MCP visibility, and a clearer, more detailed context-budget meter.
- Steadier iCloud sync across your devices.
- Tidier model settings: a faster GGUF screen, with vision options shown only for models that support them.
- Sharper math rendering for numbers and complex formulas, plus cleaner web citations.
- Fixed a model that could still appear loaded after it had been unloaded.
- Interface polish, bug fixes, and localization across every supported language.
Release 3.3MCP servers, better PDFs and datasets, evidence-backed web search, smoother cloud chats, and more model controls
Noema 3.3 brought remote Model Context Protocol servers into chat, improved how models read PDFs and work with datasets, and made cloud-model conversations steadier. Web Search gained the ability to open pages and locate the evidence behind an answer, while expanded model settings provided finer control over how each model runs. Interface fixes, bug fixes, and polish rounded out the release.
Highlights
- Connect to MCP servers: browse and add remote Model Context Protocol servers over HTTPS, then make their tools available to your models directly in chat.
- Better PDFs in chat help models read through your documents more thoroughly and find the right information faster.
- Chat turns with cloud models are smoother, with steadier and more reliable back-and-forth.
- Make fuller use of your datasets straight from chat, so answers stay grounded in your own material.
- Web Search can open pages and locate the evidence behind an answer, making it easier to inspect the sources that support a response.
- More model settings let you fine-tune how each model runs.
- Interface fixes, bug fixes, and polish throughout make the app feel more consistent and reliable.
Release 3.2Noema Autopilot, study flashcards, a reasoning toggle, faster document chats, and a refreshed Mac interface
Noema 3.2 introduced Noema Autopilot, which sizes up each message and routes it to the right model - local or remote - by difficulty, so you get the best answer while saving energy and cost. It added study flashcards built from your own documents and chats with spaced-repetition review, a new reasoning toggle to switch a model's step-by-step thinking on or off right from chat, and reworked caching that makes follow-up answers in document chats up to 2× faster - and up to 3–4× faster on large documents. A refreshed Mac interface felt cleaner and more native, alongside broad performance, stability, and interface polish throughout.
Highlights
- Turn on Noema Autopilot and every message is sized up automatically, then routed to the right model - local or remote - by how difficult it is. You get the best answer for each task while saving energy and cost.
- Study flashcards are built from your own material: turn any document or chat into a deck, then review it with spaced repetition so what you learn actually sticks.
- A new reasoning toggle lets you turn a model's step-by-step thinking on or off right from chat - deeper reasoning when you want it, faster answers when you don't.
- Reworked caching makes follow-up answers in document chats up to 2× faster - and up to 3–4× faster on large documents - so long conversations stay quick from the first reply to the last.
- A refreshed Mac interface feels cleaner, more responsive, and more at home on macOS, with layout and interactions tuned for the desktop.
- More performance, stability, and interface polish throughout, with a broad round of refinements across the app.
Release 3.1Noema 2B, a new Tools page, hands-free voice, a Logit & Jacobian Space viewer, and smoother everything
Noema 3.1 introduced our own on-device model, Noema 2B, alongside a new Tools page for benchmarking and boarding-pass scanning, a hands-free voice mode, and a new Logit and Jacobian Space viewer with vector steering and swapping on Mac. Explore gained a Trending this week section and collapsible Hugging Face READMEs in model detail views, the dataset explorer added new knowledge packs, and animations and the search bar became noticeably smoother throughout.
Highlights
- Noema 2B is our own on-device model, fine-tuned from Qwen3.5 2B with noticeably stronger performance on math, instruction following, and coding - all while staying small enough to run comfortably on-device.
- A new Tools page brings benchmarking and boarding-pass scanning together in one place, making it easier to measure performance and scan passes on-device.
- Hands-free voice mode lets you talk to Noema and hear responses back without touching your device, for a fully conversational, on-device experience. (Beta)
- A new Logit and Jacobian Space viewer lets you explore model internals with vector steering and swapping, giving you hands-on control over how the model represents and shifts meaning. (Noema Mac only)
- Model detail views in Explore now include a collapsible Hugging Face README, so you can read full model documentation without leaving the app.
- Explore adds a Trending this week section that surfaces the models people are picking up right now.
- The dataset Explore page adds new knowledge packs - curated, ready-to-use collections you can attach to your models in a tap.
- Smoother animations across the app, plus a search bar that now appears and animates cleanly instead of stuttering.
Release 3.0Faster, more native, and more reliable - smarter context, a cleaner RAG pipeline, new voice input, and a refined chat experience
Noema 3 is a major update focused on making Noema faster, more native, more reliable, and more useful across everyday tasks. It introduces improved context handling and retrieval that respects the context limit you set in Model Settings, a cleaner RAG pipeline, new on-device voice input, new pass scanning, a heavily refined chat experience, and a broad round of dataset and interface fixes. It also adds early support for Siri AI App Intents and a new CoreAI runtime for iOS 27 Beta users.
Highlights
- Improved context handling with better rolling windows and smarter truncation, plus retrieval that now respects the context limit you select in Model Settings.
- RAG has been cleaned up with clearer citations, improved evidence handling, and simplified dataset behavior.
- Early support for Siri AI App Intents lets system-level actions connect more naturally with the app. (iOS 27 Beta only)
- Voice input is a new feature, adding on-device ASR (automatic speech recognition) built on WhisperKit, with clear logging around transcription failures and a clean, simplified transcription interface.
- CoreAI support is new, bringing a dedicated runtime path, improved model loading and downloads, and fixes for unnecessary Apple Foundation Models guardrail messaging. (iOS 27 Beta only)
- The chat experience has been heavily refined - easier to scroll while a response is running, faster during downloads and long generations, and decluttered by removing redundant controls, hidden menus, and follow-up clutter.
- Dataset and enterprise handling has improved, and memory instructions are now clarified to feel optional rather than required.
- Pass scanning is a new feature, with on-device boarding-pass detection, automatic pass configuration, clear scan instructions, reliable pass-type detection, and the ability to manage and delete saved passes.
Fixes
- Many smaller fixes across model loading, bypass-RAM behavior, message action buttons, context usage indicators, storage reporting, long model-name layouts, and general interface polish.
Release 2.2Gemma 4 support, stronger imports and downloads, smarter retrieval, and chat workflow polish
Release 2.2 improves model compatibility, import reliability, document retrieval, and in-chat transparency. This update adds Gemma 4 support through a refreshed llama.cpp core, strengthens GGUF import detection, makes CML downloads easier to follow, improves how large-context models use PDFs and long documents, expands memory and system prompt customization, and ships a broad round of chat and settings fixes.
Highlights
- llama.cpp has been updated for Gemma 4 support, including fixes for previously known Gemma 4 issues.
- GGUF import is more reliable, with better detection for chat templates, JSON configs, and multimodal projector files.
- CML model downloads now show clearer progress so it is easier to understand what is happening while a model is being fetched.
- Smart retrieval makes better use of available context for PDFs and long documents, especially on large-context models.
- A new Prompt Processing card adds live progress feedback in chat, and stuck 0% progress and post-tool placement issues have been fixed.
- Memory and system prompt customization are now supported, giving you more control over how Noema behaves.
- Curated models have been refreshed with Gemma 4 and Qwen 3 1.7B support.
- Model Settings scrolling, VRAM estimates, and maximum context recommendations have been corrected to better reflect current memory-fit and KV cache quantization behavior.
Release 2.1Python tooling, OpenRouter, stronger Apple model support, and LM Studio REST v1 compatibility
Release 2.1 delivers major improvements across tooling, model support, search, Apple device compatibility, backend reliability, and chat polish. This update adds the Python tool, OpenRouter support, better Explore recommendations, direct LM Studio model downloads through the REST v1 flow, new Apple-friendly runtime support, refreshed llama.cpp and MLX integrations, and a broad round of fixes.
Highlights
- Python tool support is now built in, expanding agent workflows and local automation options.
- OpenRouter support lets you connect a broader range of hosted models from inside Noema.
- Explore recommendations have been improved to make model discovery faster and more relevant.
- LM Studio compatibility now targets REST v1, including a flow to open the Remote Endpoint button, browse Explore, select a model, and download it directly into LM Studio.
- The chat UI has been cleaned up for a more consistent, polished conversation experience.
- Executorch, CoreML, and Apple Foundation model support expand compatibility across Apple-focused runtimes.
- llama.cpp and MLX have both been updated to newer versions, including support needed for Qwen 3.5.
- Many bugs were fixed alongside backend and reliability upgrades throughout the app.
Release 2.0Native MacOS & VisionOS support, Noema Relay, Vision capabilities, and global localization
Release 2.0 brought native support for MacOS and VisionOS, introduced Noema Relay for seamless iPhone-to-Mac connection via CloudKit, added vision support for models, enhanced tool calling capabilities, rolled out localization across 10 languages, and shipped accessibility fixes.
Highlights
- Native MacOS support brings the full Noema experience to your desktop.
- VisionOS support allows you to use Noema in spatial computing environments.
- Noema Relay connects your iPhone to your Mac via CloudKit without requiring local Wi-Fi.
- Vision support for models enables photo uploads and multimodal interactions.
- Enhanced tool calling capabilities make agentic workflows more capable and reliable.
- Localization upgrades translate the Noema experience into 10 languages for global teams.
Fixes
- Accessibility fixes improved keyboard focus states, ARIA labels, and color contrast throughout the app.
- General fixes and quality-of-life updates delivered a smoother overall experience.
Release 1.4Noema is fully free with smarter performance tools
Release 1.4 makes every feature of Noema completely free while introducing FlashAttention, V cache quantization, MoE versus dense model detection, benchmarking, refreshed UI, and an updated llama.cpp core.
Highlights
- Noema is now completely free - unlimited web search and every feature are included with no upsell.
- FlashAttention and V cache quantization dramatically cut memory usage so larger models fit on more devices.
- Model catalog now distinguishes Mixture-of-Experts and dense architectures for clearer deployment decisions.
- Benchmark any model or optimization setup to compare prompt processing speed and token throughput.
- Revamped input fields and remote endpoint forms deliver a clearer, more modern interface.
- Updated llama.cpp underpinnings keep compatibility with the latest upstream improvements.
Release 1.3Remote endpoints connect your AI everywhere
Release 1.3 let you register remote inference servers alongside local models, spotlighting rich provider metadata while delivering a round of polish and fixes.
Highlights
- Remote endpoints arrive so you can connect OpenAI API, LM Studio, and Ollama servers over HTTP without leaving Noema.
- Detailed remote endpoint views highlight provider metadata including model architecture, quantization, and availability badges.
- Guided setup walks you through remote backend registration, authentication, and catalog refresh controls with smart warnings.
Fixes
- Bug fixes and reliability improvements keep hybrid chats responsive across local and remote models.
Release 1.2Reliability tune-ups and chat telemetry upgrades
Release 1.2 kept Noema responsive and insightful during chat sessions while paving the way for upcoming tool automation.
Highlights
- Improvements to token counting and speed stats in chat.
- Improved copy-pasting of chat messages for smoother sharing and note-taking.
- Added a formatted/raw toggle to the web search results popup so tool output is easier to understand.
- Laid the groundwork for a new tool call coming later these weeks (has not been updated on our github yet).
Fixes
- Bug fixes and stability improvements.
Release 1.1Stability and compatibility refinements
This update sharpens Noema's on-device reliability while preparing the app for the latest hardware. Dive into what changed below.
Fixes
- MLX model tool calling bug
- GGUF model loading bug on older devices
New features
- Added support for new iPhone 17 models.
- Updated llama.cpp framework.
- Improved overall user friendliness across the interface.
- Introduced a more in-depth onboarding process tailored for beginners.

