Open & local AI

Google launches offline AI meeting assistant with local Gemma models

Google released Google AI Edge Foresight, a Mac app that uses on-device Gemma models for offline meeting notes and document Q&A.

MacBook showing split-screen meeting notes with abstract local AI data visuals
Illustration generated for this article

Google has released Google AI Edge Foresight, a new Mac application designed to compete with AI note-takers like Granola. Launched by the same team behind an earlier experimental dictation tool, this app leverages local large language models to function entirely without an internet connection. The release marks a significant step in Google’s strategy to showcase the capabilities of its Gemma series for on-device intelligence.

What happened

The new application, Google AI Edge Foresight, is optimized specifically for Apple Silicon hardware. It utilizes the EmbeddingGemma 2 model, which contains 740 million parameters, to process and capture meeting notes across various applications. This includes support for in-person meetings, allowing users to generate transcripts and summaries locally on their machines. The app supports a split-screen interface where users can write shorthand manual notes on one side while viewing AI-generated content on the other.

Beyond real-time transcription, the tool integrates a knowledge base feature. Users can upload a variety of document formats, including PDFs, Google Docs, Microsoft Office files, plain text, Markdown, and web bookmarks. The system indexes these materials to provide context-aware answers during meetings. If a discussion point relates to data within the uploaded knowledge base, the app can answer questions in real time using a chat assistant powered by the Gemma 4 model.

This launch occurs in a increasingly crowded market for AI productivity tools. In recent months, companies such as Wispr and Calendly have introduced their own note-taking solutions. Additionally, Superhuman, formerly known as Grammarly, acquired the YC-backed note-taker Fathom last month. While Google’s previous dictation experiment did not achieve mainstream adoption, industry observers suggest this release may serve primarily to demonstrate the technical viability of offline Gemma models rather than to immediately dominate the consumer market.

How it works

The core mechanism of Google AI Edge Foresight is its reliance on local inference. By running the EmbeddingGemma 2 model directly on the device’s neural engine, the app eliminates the need to send audio or text data to cloud servers. This architecture ensures privacy and enables functionality in environments without internet access, such as during air travel or in secure facilities. The optimization for Apple Silicon allows the 740-million-parameter model to run efficiently without draining excessive battery or causing significant latency.

The knowledge base component works by embedding uploaded documents into a local vector space. When a user asks a question via the Gemma 4-powered chat interface, the system retrieves relevant information from both the live meeting transcript and the stored documents. This retrieval-augmented generation approach allows the assistant to provide accurate, context-specific answers grounded in the user’s own data, rather than relying solely on the model’s pre-trained knowledge.

Key details

  • The app is named Google AI Edge Foresight and is available for Mac.
  • It runs completely offline using the on-device EmbeddingGemma 2 model with 740 million parameters.
  • The chat assistant for answering questions is powered by the Gemma 4 model.
  • Supported upload formats include PDF, Google Docs, Microsoft Office, plain text, Markdown, and web bookmarks.
  • The interface features a split-screen view for simultaneous manual shorthand and AI-generated notes.
  • The app is optimized for Apple Silicon chips to ensure efficient local processing.

Why it matters

For software engineers and technical leads, the shift toward local-first AI tools represents a critical architectural trend. Running models on-device reduces latency and removes dependency on network stability, which is crucial for reliable productivity tools. It also addresses growing concerns about data privacy, as sensitive meeting transcripts and proprietary documents never leave the user’s hardware. This approach contrasts with cloud-based alternatives that require constant connectivity and raise potential compliance issues for enterprises handling confidential information.

Furthermore, this release highlights the maturing ecosystem of open and local models. By demonstrating that a 740-million-parameter model can handle complex tasks like real-time Q&A and transcription on consumer hardware, Google validates the feasibility of lightweight AI deployments. Developers building internal tools can look to this example when deciding whether to host models locally versus using API-based cloud services, particularly for use cases where data sovereignty is a priority.

What you can do

  • Evaluate your current meeting workflow to identify if offline capability would improve reliability or security.
  • Test local model performance on Apple Silicon hardware to understand the baseline for on-device AI tasks.
  • Review data privacy policies to determine if local-only processing meets your organization’s compliance requirements.
  • Explore the Gemma model family to assess if smaller, local models can replace larger cloud-based APIs for specific tasks.
  • Consider integrating local vector search capabilities into internal knowledge management systems for faster, private retrieval.
  • Monitor the competitive landscape of AI note-takers to benchmark features against emerging local-first solutions.

Tools from the Bytechap store

Keep reading

All stories