Custom Google Gemini integration development

Gemini is Google's multimodal model family: a single API that processes text, images, audio, video and PDFs together, with context windows of up to 2 million tokens and native Google Search grounding. We build production-grade Gemini integrations — via the Gemini API for rapid iteration, or via Vertex AI for enterprise-grade deployment with EU data residency. For SaaS products, Workspace augmentation, video analysis pipelines and search-grounded RAG.

Gemini 2.5 Pro Gemini 2.0 Flash Vertex AI Multimodal Search grounding Function calling EU multi-region
Discuss your Gemini case View applications
prompt · multimodal img audio video gemini-2.5-pro grounded answer [1] source [2] source function_call tool vertex-ai

Why choose Gemini over other LLMs

Gemini stands out from the OpenAI and Anthropic models in four respects: native multimodality, an exceptionally long context window, built-in Google Search grounding and direct connection to Google Workspace and Google Cloud. For certain use cases, Gemini is not an alternative but the logical choice.

The Gemini family consists of several models, each serving a different price and performance point. Gemini 2.5 Pro is the flagship, with the most demanding reasoning capabilities and a context window that can stretch to 2 million tokens, enough to feed a complete legal case file, an entire codebase or several hours of video into a single prompt. Gemini 2.0 Flash is the cost- and latency-optimised model for high throughput, ideal for chatbots, classification tasks and real-time applications where response time is critical. Gemini 1.5 Pro remains available as a stable production version. For on-device use on Android devices there is Gemini Nano, which runs locally on the Pixel 8 Pro and newer, with no network traffic and no data leaving the device.

The architectural difference from other models lies in native multimodality. Whereas GPT-4 and Claude process text and images through separate components, Gemini was trained from the ground up on text, code, images, audio and video simultaneously. This means you can supply an hour of video as input and ask the model nuanced questions about what happens at specific moments, without first extracting transcripts or frames. That same architecture makes possible the real-time applications we now see in Astra-style demos: a live audio and video stream to which Gemini responds directly.

The Google Search integration is unique. With the grounding tool, Gemini retrieves live Google search results for each answer and cites them as sources. For applications where factual accuracy and up-to-date information are critical, such as research tools, support bots and analyst assistants, this replaces a custom RAG pipeline with a vector database. Read the official Gemini API grounding documentation for the technical details.

What we are building with Gemini for our clients

The combination of multimodal input, long context and search grounding opens up applications that are difficult or expensive with other models. Here are some of the patterns we have implemented recently.

📑

Workspace augmentation and Docs, Sheets and Gmail AI

Gemini integrates directly with Google Workspace via add-ons and Apps Script. We build Docs extensions that summarise contracts, Sheets functions that classify data, and Gmail integrations that suggest replies based on the full thread history. For organisations already on Workspace, this is the quickest route to AI functionality without switching context.

🎥

Video analysis pipelines

Gemini can process an hour of video directly as input. We build pipelines that automatically transcribe, index and make searchable by event recorded support calls, production footage or security material. There is no separate Whisper transcription and vision step: a single API call returns structured output with timestamps and sentiment.

🌐

Search-grounded RAG without a vector DB

With Google Search grounding, you do not need to set up a separate vector database, embeddings pipeline and retrieval layer for every application. For research, marketing and analyst tools, we let Gemini query the web directly and cite sources. For proprietary knowledge, we combine this with Vertex AI Search or a custom RAG layer.

📞

Real-time multimodal apps (Astra-style)

With Gemini's Multimodal Live API, we build applications where users can share their camera, microphone and screen at the same time, and the model responds live. Uses include visual support assistants, instructional video coaches, and on-site inspection tools where an engineer points their phone at an installation and receives real-time feedback.

📱

On-device AI with Gemini Nano

For Android apps, Gemini Nano runs locally via AICore and the ML Kit GenAI APIs. We integrate on-device summarisation, smart reply and proofreading with no network traffic and no data leaving the device. This is crucial for privacy-sensitive sectors such as healthcare, finance and legal work.

🧮

Code execution and data analysis agents

Gemini's built-in code execution sandbox lets the model write and run Python snippets itself. We build analyst agents that read in spreadsheets, CSVs or database exports, carry out statistical analyses and return visualisations, without having to define every calculation as a separate tool.

Our approach to Gemini integrations

We build Gemini integrations that are production-ready from day one: with streaming responses, robust error handling, cost monitoring per call, prompt versioning and evals that catch regressions before they reach users. No demo-grade prototypes that have to be rebuilt for production.

For every project we start with a choice between three deployment paths. Google AI Studio and the Gemini API directly (via generativelanguage.googleapis.com) are ideal for quick proofs of concept and for consumer products where EU residency is not a hard requirement. Vertex AI is the enterprise route: the same models, but through Google Cloud with IAM integration, VPC Service Controls, audit logs in Cloud Logging, customer-managed encryption keys and, critically for Dutch organisations, a choice of data centre in europe-west4 (Eemshaven) or other EU regions, with the guarantee that your data is not used for model training. For clients in healthcare, finance and the public sector, Vertex AI is almost always the right choice.

We work with the official Google SDKs for Python, JavaScript/TypeScript, Go, Java and Swift. For production applications we use streaming responses by default (so tokens appear in the UI immediately), function calling so Gemini can interact with your existing systems, structured output via JSON mode or response schemas for type-safe integrations, and safety settings to tailor content filtering to your use case. For batch jobs we use Gemini's batch mode, which is up to 50% cheaper than synchronous calls.

From first prompt to production deployment

Our Gemini projects go through four phases. Each phase delivers a working result that you can test yourself, with no months of architecture discussions without working code.

Use-case scoping and model selection

We determine which Gemini model suits your use case: 2.5 Pro for demanding reasoning, 2.0 Flash for latency-critical tasks, Nano for on-device. We choose between the Gemini API directly and Vertex AI based on data residency, IAM and compliance requirements.

Prompt engineering and proof of concept

In Google AI Studio we build the first working prompts, including few-shot examples, system instructions and function-call schemas. Within one to two weeks there is a testable prototype that you can validate yourself.

Production implementation

The prompts move into a versioned repository, we build the SDK integration into your application (Python, Node, Go), enable streaming, and implement caching, retry logic and cost monitoring per request via Cloud Logging or your own telemetry layer.

Evals, monitoring and refinement

We build an evaluation suite that catches regressions when prompts or models change. In production we monitor latency, token usage, hallucination rates and user feedback, and adjust when a new Gemini model is released or the use case shifts.

Our technical stack for Gemini projects

The Gemini API is rich in features: function calling, code execution, JSON mode, response schemas, safety settings, system instructions, file uploads for video and PDF, batch mode and the Multimodal Live API for real-time audio and video. We deploy them deliberately: not every feature is relevant to every application, but together they form a powerful foundation.

Around the Gemini calls, we build the standard production components: a queue mechanism for batch work, a prompt registry with versioning, evaluation suites that we run against every deployment, observability via OpenTelemetry or Vertex AI's own tooling and, for organisations that want it, a feedback loop through which users can flag poor answers so that we can improve the prompts and model choice.

Gemini 2.5 Pro Gemini 2.0 Flash Gemini 1.5 Pro Gemini Nano Vertex AI Google AI Studio google-genai SDK Function calling JSON mode Code execution Search grounding Multimodal Live API Batch prediction Cloud Run Cloud Functions Apps Script Workspace Add-ons europe-west4

GDPR and data residency on Vertex AI

For Dutch organisations, particularly in healthcare, finance, government and legal work, data residency and a guarantee that prompts are not used for model training are not optional but a requirement. Vertex AI addresses this in a way that the Gemini API does not offer directly.

EU multi-region and data centre choice

On Vertex AI, you can call Gemini models from specific EU regions: europe-west4 (Eemshaven, the Netherlands), europe-west1 (Belgium), europe-west3 (Frankfurt) and the broader EU multi-region. Data at rest and data in transit remain within the EU. This is contractually documented in the Google Cloud Service Specific Terms for Vertex AI.

No training on customer data

Google explicitly confirms that Vertex AI prompts and outputs are not used for training or fine-tuning foundation models. For the free tier of the Gemini API, the opposite applies, a key difference that many teams overlook. We run sensitive workloads on Vertex AI by default.

IAM, VPC-SC and CMEK

Vertex AI integrates with Google Cloud IAM for access control to models and data, with VPC Service Controls for network isolation, and with Customer-Managed Encryption Keys (CMEK) so that you manage the keys yourself. Audit logs land in Cloud Logging and can be exported to your SIEM.

DPIA support and DPA

For processing under the GDPR, we provide the input for your DPIA: which data goes to Gemini, with what retention, and which safety settings are active. Google Cloud's Data Processing Addendum covers the processor agreements at platform level.

Concrete Gemini scenarios from projects

Patterns we have recently built or advise clients on: concrete and realistic, not as a vision of the future but as things that can run in production today.

Contract analysis across a full case file

A legal services provider processes acquisition files running to hundreds of pages. With Gemini 2.5 Pro's 2-million-token context window, complete NDAs, purchase agreements and annexes fit into a single prompt. The model identifies risk clauses, compares versions and produces a reasoned recommendation, including references to specific page numbers in the source document.

Video feedback for instructional content

A training organisation has trainers record their own instructional videos. Gemini analyses tone of voice, structure, audio quality and visual clarity and provides structured feedback per video within minutes, including timestamps for improvement points. Previously, this took an internal reviewer half a day per video.

Workspace add-on for quote generation

A SaaS company builds quotes in Google Docs. A Workspace add-on we built calls Gemini to generate a draft quote, with the correct pricing table and legal clauses, based on CRM data (product choice, customer segment, region), directly in the Doc, with the house style already applied. The account manager checks it and sends it.

Search-grounded market analysis tool

A consultancy firm receives questions about market trends, competitor moves and regulation every day. Using Gemini's Google Search grounding, we build an internal tool that answers these questions directly from the web and cites its sources. It is faster than searching manually, and it provides an audit trail of the sources used.

Why choose Appfront for your Gemini project

Multi-LLM experience since GPT-3

We have been building LLM integrations since the GPT-3 beta. We know not only Gemini but also OpenAI, Anthropic and open-source models, and we know which model works best for which task. That difference often shows up in total cost of ownership and final quality.

In-house Google Cloud expertise

Vertex AI is not a standalone product but part of the Google Cloud ecosystem. We set up deployments with IAM, VPC-SC, BigQuery integrations for logging and Cloud Run for serverless inference, all in a way that fits within your Cloud organisational policy.

Production quality, not demo prototypes

Streaming, retries, cost tracking, prompt versioning, evaluations and monitoring are part of the core implementation, not something we add afterwards. The result: solutions that carry over to new model versions and capture user feedback.

Frequently asked questions about Gemini integrations

When should I choose Gemini over OpenAI or Claude?
Three scenarios point to Gemini: you need to process video or audio directly (no other frontier model does this natively), you have extremely long documents where a context of 1 to 2 million tokens is useful, or you already run on Google Workspace and Google Cloud and want straightforward IAM integration and EU data residency via Vertex AI. For pure text tasks where none of these apply, OpenAI and Anthropic are often equivalent or better.
What is the difference between the Gemini API, AI Studio and Vertex AI?
AI Studio is the browser-based IDE for quickly testing prompts, which is useful for proofs of concept. The Gemini API (generativelanguage.googleapis.com) is the direct API endpoint for consumer and SaaS applications, with a free tier. Vertex AI is the enterprise route via Google Cloud: the same models, but with IAM, VPC Service Controls, customer-managed encryption keys, EU data centres, audit logs and a guarantee that data is not used for training. For business production workloads, we almost always choose Vertex AI.
Can Gemini be used GDPR-compliantly for Dutch customer data?
Yes, provided you use Vertex AI with EU regions (such as europe-west4 in Eemshaven). On Vertex AI, prompts and outputs are not used for model training, data residency is contractually guaranteed, and the Google Cloud Data Processing Addendum covers the processor agreements. The free tier of the Gemini API does not offer these guarantees, as data there may be used for training.
How much text, audio or video can I include in a single Gemini prompt?
Gemini 2.5 Pro supports context windows of up to 2 million tokens, equivalent to roughly 1,500 pages of text, 22 hours of audio or 2 hours of video. Gemini 2.0 Flash offers 1 million tokens. That is considerably more than current OpenAI and Anthropic models. In practice, we use long contexts mainly for legal case files, codebase analysis and long-form video content.
How exactly does Google Search grounding work?
You enable the grounding tool as a parameter in the API call. Gemini decides for itself when a Google search is needed, runs it, reads the results and uses them to build an answer, with explicit references to the sources used (URLs). For applications where factual accuracy is critical, this can replace a custom RAG pipeline. The feature is available on both the Gemini API and Vertex AI.
What is function calling and when should I use it?
Function calling lets Gemini indicate in a structured way that it wants to call an external tool, such as your database, an ERP API or an external service. You define the functions (name, parameters, return schema), and the model decides for itself when and how to call them. Essential for agent-style applications where the model needs to interact with the outside world without us writing a separate prompt for every action.
What about the cost of Gemini in production?
Gemini is in the same range as OpenAI's GPT models, with Flash variants that are considerably cheaper than Pro variants. Batch mode offers up to 50% off synchronous pricing. Cost monitoring is something we build in as standard, per request, per user and per feature, so you can quickly see which prompts are getting out of hand. Current prices are on Google's official AI pricing page and change from time to time.
Can I use Gemini Nano on-device in my Android app?
Yes, for Android 14+ on compatible devices (Pixel 8 Pro, Samsung Galaxy S24 and newer), Gemini Nano runs locally via AICore and the ML Kit GenAI APIs. We integrate this for summarisation, smart reply and proofreading without data ever leaving the device. On unsupported devices, the app falls back to the cloud API or another on-device strategy.

Ready to build Gemini into your product?

Discuss your use case with us. Together we'll look at whether Gemini is the right choice, which deployment route suits your requirements and what a first working prototype could look like, with no obligation.

Schedule a conversation

Edit content