> ## Documentation Index
> Fetch the complete documentation index at: https://docs.lovable.dev/llms.txt
> Use this file to discover all available pages before exploring further.

# AI features for your app

> Add AI features like chatbots, summaries, and image generation to your Lovable app with the built-in AI connector. No API keys or provider setup required.

Normally, adding AI to an app means choosing a provider, managing API keys, setting up billing, securing credentials, and wiring model calls into your backend. Lovable’s built-in AI connector handles that setup for you, so you can add AI features to your app by describing what you want to build.

These AI features run inside your app. They are separate from the Lovable agent that helps you build and edit your project.

Some examples include:

* **Summaries**: automatically condense long text into clear takeaways.
* **Chatbots and assistants**: build conversational helpers into your app.
* **Sentiment detection**: understand user feedback at scale.
* **Document Q\&A**: let users ask questions directly against your content.
* **Creative generation**: brainstorm ideas, draft copy, or expand concepts.
* **Translation**: serve users across languages.
* **Image and document analysis**: extract, summarize, and interpret key information from unstructured content.
* **Workflow automation**: handle repetitive or multi-step tasks inside your app.
* **Semantic search and retrieval-augmented generation (RAG)**: search documents, knowledge bases, and content by meaning instead of exact keywords.
* **Text-to-speech**: turn text into spoken audio for voice narration, read-aloud, and audio-first experiences.
* **Speech-to-text**: transcribe voice notes, recordings, and meetings, and add voice input or dictation.
* **Live voice conversations**: let people talk to your app, interrupt it, and have it act on what they say.
* **Music generation**: generate short clips or full songs, with vocals or as instrumentals, for music makers, jingles, and soundtracks.
* **Classification and routing**: assign tickets, flag policy issues, rank candidates, or score quality with typed decisions instead of free-form chat.

For copy-paste prompts covering several of these features, see the [AI section of the prompt library](/prompting/prompting-library#ai-features).

## Enabling the built-in AI connector

<Note>
  For the best experience, use the built-in AI connector together with the [built-in backend (Cloud)](/features/cloud), so your app can make secure model calls.
</Note>

By default, the built-in AI connector is enabled for your workspace, and Lovable can add AI features to your app when requested. To manage the built-in AI connector for your projects, open **Connectors**, select **AI**, and adjust the settings under **Manage my agent's permissions**.

Workspace admins can also disable the built-in AI connector entirely for the workspace from **Connectors** → **Admin settings**.

### Permission preferences

The default setting is **Always allow**, meaning the built-in AI connector can be used automatically in your projects. You can change your preference anytime: open **Connectors**, select **AI**, and adjust the settings under **Manage my agent's permissions**.

Choose between:

* **Always allow**: Lovable automatically performs the action without asking for review or approval.
* **Ask each time**: Lovable asks for your approval whenever the action is needed. For example, if you want to add a chatbot, you can:
  * **Allow**: enable the integration for the current project.
  * **Deny**: decline the integration for this request. You may be asked again later.
  * **Adjust preferences**: change the default behavior for future projects. This does not affect the current project.
* **Never allow**: Lovable blocks the action, informs you that AI is required, and instructs you to enable the built-in AI connector.

## How it works

Lovable sets up the AI infrastructure for you:

* **API key**: Lovable automatically generates and manages a `LOVABLE_API_KEY` for each project. You never need to create or provide it yourself. When a project is remixed, a fresh key is generated for the new project automatically.
* **Backend calls**: AI calls run through a secure backend function that Lovable creates for you, which keeps your credentials and prompts server-side. The exception is live voice: the browser streams audio directly to the voice model, while your app's server still starts the call and keeps the API key.
* **Streaming**: the built-in AI connector supports streaming responses with server-sent events (SSE). Lovable uses streaming by default for chatbot and assistant features, so responses appear token by token rather than all at once.

## Supported models for AI features in your app

<Note>
  These models are available for AI features inside the apps you build, such as chatbots, semantic search, voice, typed decisions, and image, video, and music generation. They are not the models Lovable uses to write, edit, or reason about your code.
</Note>

When you ask Lovable to add or update an AI feature, you can name a supported model or describe what you want and let Lovable pick the right one. Each model links to its official source for technical details. To browse the models available to your workspace in the product, open **Connectors**, select **AI**, and search the **Available models** list by model name, provider, or capability.

Models marked **(deprecated)** keep working in apps that already use them, but Lovable does not use them when it builds new AI features. If you ask for a deprecated model, Lovable recommends a supported model instead. In **Connectors** → **AI**, deprecated models show a **Deprecated** badge.

<Note>
  Enterprise workspaces use only zero-data-retention models unless a workspace admin enables [Extended-retention models](/features/privacy-and-security-settings#extended-retention-models). The Claude models and Gemini Omni 1.1 Flash are not zero-data-retention models, so on Enterprise plans your app can use them only after an admin enables that setting. When an admin enables [EU inference](/features/eu-inference), your app can use only the models that run in the European Union. A request for any other model fails with an error.
</Note>

### Choosing a model

Not sure which to use? Describe what you want, and Lovable picks a model for you. This table shows where to start and when you might switch.

| Use case | Start with | Switch when |
| - | - | - |
| Chat and assistant features | Gemini 3.8 Flash | You need deeper reasoning or longer context. |
| High-volume, simple text tasks | Gemini 3.1 Flash Lite or GPT-5 Nano | Accuracy matters more than cost. |
| Image generation and editing | GPT Image 2 | You want faster results (GPT Image 2.5 Flare), the highest image quality (GPT Image 2.5 Sunburst), lower-cost drafts, or a Gemini image model. |
| Video generation | Gemini Omni 1.1 Flash | You want the highest visual quality (Veo 3.1), or you need the same prompt to produce a near-identical clip each time (any Veo model). |
| Music generation | Lyria 3 Clip Preview | You want a full track longer than 30 seconds (Lyria 3 Pro Preview). |
| Semantic search and retrieval-augmented generation (RAG) | Gemini Embedding 2 | You want lower cost on high-volume text workloads. |
| Classification, routing, and scoring | Jev Latest | You need generated prose, explanations, or extraction of new text. |
| Text-to-speech | GPT-4o Mini TTS | You need character or higher-fidelity voices. |
| Speech-to-text | Gemini 3.5 Transcribe | Your audio files are larger than 14 MB. |
| Live voice conversations | GPT-Live 1 | You only need to read text aloud or transcribe audio, which the text-to-speech and speech-to-text models handle at lower cost. |

### Chat models

Chat models power conversational and text features: chatbots, assistants, summaries, document Q\&A, translation, classification, and extraction. They range from fast, low-cost models for simple tasks to high-reasoning models for complex work.

Ask Lovable to:

* Build a chatbot or in-app assistant.
* Summarize long text, documents, or transcripts.
* Answer questions from your own content.
* Classify, extract, or translate text.

| Model | Best for | Priority processing |
| - | - | - |
| [Gemini 3.8 Flash (default)](https://deepmind.google/models/model-cards/gemini-3-8-flash/) | Latest Flash generation for fast coding, reasoning, and agentic workflows. Accepts text, image, audio, and video input. | Yes |
| [Gemini 3.7 Flash](https://deepmind.google/models/model-cards/gemini-3-7-flash/) | Fast coding, reasoning, and agentic workflows. Accepts text, image, audio, and video input. | Yes |
| [Gemini 3.6 Flash (deprecated, stops working on November 19, 2026)](https://deepmind.google/models/model-cards/gemini-3-6-flash/) | Fast coding, reasoning, and agentic workflows. Accepts text, image, audio, and video input. | Yes |
| [Gemini 3.5 Flash (deprecated)](https://deepmind.google/models/model-cards/gemini-3-5-flash/) | Fast coding, reasoning, and agentic workflows. | Yes |
| [Gemini 3.1 Pro Preview](https://ai.google.dev/gemini-api/docs/models/gemini-3.1-pro-preview) | Advanced coding, long-context understanding, and complex multi-step reasoning. Slower and premium-priced. | Yes |
| [Gemini 3.1 Flash Lite](https://ai.google.dev/gemini-api/docs/models/gemini-3.1-flash-lite) | High-volume, lightweight tasks like classification, summarization, and translation. | Yes |
| [Gemini 3 Flash Preview](https://ai.google.dev/gemini-api/docs/models/gemini-3-flash-preview) | Fast, general-purpose chat and iterative builds where responsiveness matters. | Yes |
| [Gemini 2.5 Pro (deprecated)](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/2-5-pro) | Deep reasoning, advanced coding, and research. Most capable 2.5 model, most expensive. | Yes |
| [Gemini 2.5 Flash (deprecated)](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/2-5-flash) | Assistants and general workflows balancing speed and intelligence. | Yes |
| [Gemini 2.5 Flash Lite (deprecated)](https://docs.cloud.google.com/vertex-ai/generative-ai/docs/models/gemini/2-5-flash-lite) | Simple, high-throughput tasks at the lowest cost. | Yes |
| [GPT-6 Astra](https://platform.openai.com/docs/models/gpt-6-astra) | OpenAI's most capable model for complex reasoning, coding, research, and document creation. Premium-priced. | Yes |
| [GPT-6 Sol](https://platform.openai.com/docs/models/gpt-6-sol) | Mid-tier GPT-6 model for complex coding and agentic workflows at lower cost than GPT-6 Astra. | Yes |
| [GPT-6 Luna](https://platform.openai.com/docs/models/gpt-6-luna) | Low-cost GPT-6 model for simple or high-volume tasks. | Yes |
| [GPT-5.6 Sol](https://platform.openai.com/docs/models/gpt-5.6-sol) | Flagship GPT-5.6 preview. Strongest for the hardest reasoning, coding, and agentic tasks. | Yes |
| [GPT-5.6 Terra](https://platform.openai.com/docs/models/gpt-5.6-terra) | Balanced GPT-5.6 preview for everyday work at lower cost. | Yes |
| [GPT-5.6 Luna](https://platform.openai.com/docs/models/gpt-5.6-luna) | Fast, low-cost GPT-5.6 preview for simple or high-volume tasks. | Yes |
| [GPT-5.5 Pro](https://platform.openai.com/docs/models/gpt-5.5-pro) | Frontier reasoning, in-depth research, and complex engineering. Slowest and most expensive. [Always reasons before answering](#show-the-models-thinking), so it can't be used for plain chat replies. | No |
| [GPT-5.5](https://platform.openai.com/docs/models/gpt-5.5) | Complex reasoning, advanced coding, and long-context knowledge work. | Yes |
| [GPT-5.4 Pro (deprecated)](https://platform.openai.com/docs/models/gpt-5.4-pro) | Advanced coding, deep research, and long-context multi-step reasoning. Premium-priced. [Always reasons before answering](#show-the-models-thinking), so it can't be used for plain chat replies. | No |
| [GPT-5.4](https://platform.openai.com/docs/models/gpt-5.4) | Complex reasoning, coding, and long-context knowledge tasks. | Yes |
| [GPT-5.4 Mini](https://platform.openai.com/docs/models/gpt-5.4-mini) | Assistants and mid-complexity reasoning at lower cost. | Yes |
| [GPT-5.4 Nano](https://platform.openai.com/docs/models/gpt-5.4-nano) | Summaries, classification, and high-volume simple tasks. Cheapest and fastest 5.4. | No |
| [GPT-5.2 (deprecated)](https://openai.com/index/introducing-gpt-5-2/) | Complex reasoning and deep coding or analytical workflows. | Yes |
| [GPT-5 (deprecated)](https://openai.com/index/introducing-gpt-5/) | Accuracy-critical tasks and high-quality reasoning. | Yes |
| [GPT-5 Mini (deprecated)](https://platform.openai.com/docs/models/gpt-5-mini) | Assistants and business workflows balancing speed and cost. | Yes |
| [GPT-5 Nano](https://platform.openai.com/docs/models/gpt-5-nano) | Quick, simple responses and high-volume tasks. Cheapest and fastest GPT-5. | No |
| [Chat Latest](https://developers.openai.com/api/docs/models/chat-latest) | The latest Instant model ChatGPT uses, tuned for conversational chat. Accepts text and image input. It doesn't reason before answering, so it can't show its thinking. OpenAI regularly changes the model behind this name, and what it costs can change with it, so pick a numbered GPT model when you need responses and cost to stay predictable. | No |
| [Claude Fable 5.1](https://platform.claude.com/docs/en/models/fable-5-1/overview) | Anthropic's most capable model for demanding reasoning and long-running, multi-step tasks. Premium-priced. | No |
| [Claude Opus 5.5](https://platform.claude.com/docs/en/models/opus-5-5/overview) | Long-running, multi-step coding and knowledge work. | Yes |
| [Claude Opus 5](https://platform.claude.com/docs/en/models/opus-5/overview) | Complex coding and multi-step tasks. | Yes |
| [Claude Sonnet 5](https://platform.claude.com/docs/en/models/sonnet-5/overview) | Balanced speed and intelligence for everyday chat and text work. | No |
| [Claude Haiku 4.5](https://platform.claude.com/docs/en/models/haiku-4-5/overview) | Anthropic's fastest model for straightforward, high-volume tasks. | No |

The Claude models accept text, image, and PDF input. Lovable adds them to newer apps (created from May 13, 2026, which use [TanStack Start](/features/upgrade-to-tanstack-start)). Older React + Vite apps use the other chat models.

#### Show the model's thinking

Many models work through a problem before they answer. Your app can show that thinking above the answer as it streams in, so people see progress during a long answer instead of waiting at a blank screen. What appears is a summary of how the model worked through the problem: the raw chain of thought is never exposed, and nothing appears unless your app asks for it.

Ask Lovable to:

* Show the model's thinking while it answers.
* Stream the thinking into a chat interface above the reply.

This works on the model your app already uses, so you don't need to switch models to get it. Most chat models in the table above can do it.

Thinking counts toward the answer's output, so an answer that shows its thinking costs more than the same answer without it. Answers also take longer, sometimes minutes for the most capable models. See [Usage and pricing](#usage-and-pricing).

<Note>
  **GPT-5.5 Pro** and **GPT-5.4 Pro** work differently from the other chat models: they always reason before answering, and they can't be used for plain chat replies. They are the slowest and most expensive models available, so use them only when a task really needs that depth.
</Note>

#### Faster responses with priority processing

If low latency matters, you can ask Lovable to use **priority processing** for a chat feature. Priority processing sends the request to the model provider's faster serving tier, so responses can come back more quickly during busy periods.

Ask Lovable to:

* Make an AI feature respond faster or with lower latency.
* Use priority processing for a chat feature.

Priority processing is available for supported chat models only. The **Priority processing** column in the [chat model table](#chat-models) shows which models support it, and asking for it on any other model has no effect. Priority processing comes at a premium price: you are charged the priority rate only when the provider actually serves the request at the priority tier. If the provider serves the request at the standard tier instead, you pay the standard rate. See [Usage and pricing](#usage-and-pricing) for how AI usage is billed.

Priority processing applies to chat requests only. Image, video, embedding, voice, and typed-decision models do not support it.

### Image models

Image models generate and edit images from text prompts, uploaded images, or both. Use them for visual assets, product mockups, marketing imagery, and in-app image editing.

Ask Lovable to:

* Generate images from a text description.
* Edit or restyle an uploaded image.
* Create product mockups or marketing visuals.
* Produce thumbnails or illustrations on demand.

| Model | Best for |
| - | - |
| [GPT Image 2 (default)](https://platform.openai.com/docs/models/gpt-image-2) | High-quality image generation and editing for product mockups, marketing imagery, and creative content. |
| [GPT Image 2.5 Flare](https://platform.openai.com/docs/models/gpt-image-2.5-flare) | Fast image generation and editing when response time matters most, such as previews and in-app image editors. |
| [GPT Image 2.5 Sunburst](https://platform.openai.com/docs/models/gpt-image-2.5-sunburst) | Highest-quality image generation and editing when output quality or editing precision matters most, such as final assets and detailed visuals. |
| [GPT Image 1 Mini (deprecated)](https://platform.openai.com/docs/models/gpt-image-1-mini) | Cost-efficient image generation and editing for thumbnails, drafts, and high-volume workflows. |
| [Gemini 3.1 Flash Image](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-1-flash-image) | Fast image generation and editing with strong subject consistency and text rendering. Also known as Nano Banana 2. |
| [Gemini 3.1 Flash Lite Image](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-1-flash-lite-image) | Faster, lower-cost image generation and editing for high-volume, cost-sensitive workflows. Also known as Nano Banana 2 Lite. |
| [Gemini 3 Pro Image](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/3-pro-image) | Detailed visuals, text rendering in images, and multi-image composition. Also known as Nano Banana Pro. |
| [Gemini 2.5 Flash Image (deprecated)](https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini/2-5-flash-image) | Very low-cost image generation and quick visual outputs. |

All the image models above can edit an uploaded image as well as generate new ones. Describe the change you want, and with the OpenAI models you can also limit the edit to part of the image with a mask. GPT Image 2.5 Flare and GPT Image 2.5 Sunburst can produce images with a transparent background, as PNG or WebP files.

When you ask Lovable to add image generation to your app, you can ask for a speed or quality level instead of a specific model: fast, standard, or premium. Lovable picks the model for each level and updates it as better models become available. Fast currently uses GPT Image 2.5 Flare, and standard and premium use GPT Image 2.5 Sunburst.

### Video models

Video models generate short video clips with sound. You describe the scene in a text prompt, or provide an image for the model to bring to life. Use them for product teasers, animated social posts, background loops, and storyboard previews.

Ask Lovable to:

* Generate a video clip from a text description.
* Animate an uploaded image into a short clip.
* Produce a vertical clip for social media.
* Feature a specific product, person, or character in a new scene by supplying reference photos of the subject.
* Turn two images into a clip that starts on one and ends on the other.
* Change what an existing video shows by describing the edit, whether the clip was generated earlier or uploaded by your app's users.
* Extend an existing video to build a longer one.

| Model | Best for |
| - | - |
| [Veo 3.1 Lite (default)](https://ai.google.dev/gemini-api/docs/video) | Low-cost clips for drafts, social posts, and high-volume use. Doesn't support 4K. |
| [Veo 3.1 Fast](https://ai.google.dev/gemini-api/docs/video) | Better quality at moderate cost, with 4K support. |
| [Veo 3.1](https://ai.google.dev/gemini-api/docs/video) | The highest quality and by far the most expensive. |
| [Gemini Omni 1.1 Flash](https://ai.google.dev/gemini-api/docs/omni) | The latest generation: fast, low-cost clips from 360p drafts up to 4K, plus editing and extending existing videos, including uploaded footage. |

A new clip arrives as an MP4 file with sound, in landscape (16:9) or portrait (9:16). The models differ in clip length, resolutions, and what they can do with existing footage:

| | Gemini Omni 1.1 Flash | Veo models |
| - | - | - |
| Clip length | 3 to 10 seconds, at any resolution. | 4, 6, or 8 seconds. 1080p and 4K clips are always 8 seconds. |
| Resolutions | 360p, 720p, 1080p, and 4K. The 1080p and 4K outputs are upscaled from a lower resolution. | 720p, 1080p, and 4K. 4K needs Veo 3.1 Fast or Veo 3.1. |
| Reference photos | Any clip length. Can be combined with other inputs. | Needs Veo 3.1 or Veo 3.1 Fast, 1 to 3 photos, and an 8-second clip. Can't be combined with other inputs. |
| Editing | Videos up to 10 seconds long and about 45 MB, generated or uploaded. | Not supported. |
| Extension | Adds 3 to 10 seconds to the end of the clip per extension, up to about 40 seconds in total. Works on generated and uploaded videos up to about 45 MB. | Adds 7 seconds per extension to a clip the model generated. |
| Input billing | Attached images and clips add a small charge, billed as input tokens. | Inputs are not charged. |

On Gemini Omni 1.1 Flash, a 360p clip costs about a third of the same clip at 720p, which makes 360p a low-cost way to draft a clip before regenerating it at full quality.

Instead of describing everything in text, you can start from your own images:

* **Animate an uploaded image.** The clip starts from your exact photo and sets it in motion, so the photo's scene and framing stay.
* **Supply two images as the first and last frame.** The model generates the motion between them, which is a simple way to make a transition between two visuals. This works on Gemini Omni 1.1 Flash, Veo 3.1, and Veo 3.1 Fast.
* **Supply reference photos of a subject.** Provide photos of the same product, person, or character, and the model places that subject into a freshly generated scene described by your prompt. Unlike animating a photo, nothing from the photos appears directly in the clip: the photos only fix what the subject looks like, and the scene around it is newly generated. Several angles of the same subject improve fidelity.

Gemini Omni 1.1 Flash also works with existing footage, whether Lovable generated it earlier or your app's users uploaded it. To edit a video, describe one change, such as swapping the background or replacing an object, and Lovable regenerates the clip with that change applied at the same length. To build a longer video, ask Lovable to extend a clip: each extension continues the same scene and produces a single merged video, which you can extend again. The Veo models can also extend clips they generated. Either way, only the new seconds are billed.

Generating a clip usually takes around 1-2 minutes, so video works differently from chat or images: the clip is generated in the background, and your app checks on it until it is done. Write prompts in English. The Veo models require it, and Google recommends it on Gemini Omni 1.1 Flash. The Veo models also rewrite every prompt before generating, so the finished clip follows the intent of your prompt rather than its exact wording.

A finished clip stays available to your app for **48 hours**. If your app should show a video again later, or keep it permanently, ask Lovable to save it to your app's own storage.

Generating video uses credits, billed per second of generated video. Higher resolutions cost more per second, and a single clip costs far more than a typical chat response. You are charged when a generation completes, even if your app stops waiting or never shows the clip. Failed generations are not charged. Each project also generates a limited number of clips at the same time: one on Free workspaces, and up to ten on paid plans. See [Usage and pricing](#usage-and-pricing) and [Workspace rate limits](#workspace-rate-limits).

### Music models

Music models generate music from a text prompt, with vocals or as an instrumental. You can also attach an image for inspiration. The image and the prompt together must stay under 1 MB. Use music models for music makers, soundtracks, jingles, and background music in your app.

Ask Lovable to:

* Build a music maker where people describe a track and play the result.
* Add a soundtrack to a story, lesson, or product page.
* Generate an instrumental background track in a style you describe.
* Turn lyrics you provide into a song.

| Model | Best for |
| - | - |
| [Lyria 3 Clip Preview](https://ai.google.dev/gemini-api/docs/models/lyria-3-clip-preview) | Clips of about 30 seconds with vocals or instrumentals, for previews, jingles, and short loops. |
| [Lyria 3 Pro Preview](https://ai.google.dev/gemini-api/docs/models/lyria-3-pro-preview) | Full tracks up to 184 seconds with vocals or instrumentals. |

Describe the track in the prompt: genre, mood, tempo, instruments, and singing language. To use your own lyrics, include them in the prompt. For a full track, you can also describe the sections in order, such as an intro, a verse, and a chorus.

The length depends on the model. Lyria 3 Clip Preview produces clips of about 30 seconds. Lyria 3 Pro Preview produces tracks up to 184 seconds, and you can ask for a length in the prompt to guide it.

A finished track arrives as one MP3 file when the whole track is generated, so your app shows a waiting message until then. Lovable does not store generated tracks. To let people play a track again later, ask Lovable to save it to your app's own storage.

Generating music uses credits, charged per track when it completes, even if your app stops waiting for it. Failed generations are not charged. Music models are not available while [EU inference](/features/eu-inference) is enabled.

### Embedding models

Embedding models turn content into a format that can be searched by meaning instead of exact keywords. Use them for semantic search, retrieval-augmented generation (RAG), FAQ bots, document search, and knowledge bases.

Ask Lovable to:

* Build semantic search over uploaded documents.
* Create a FAQ bot that answers from your help content.
* Build a company knowledge base that finds relevant internal docs.

| Model | Best for |
| - | - |
| [Gemini Embedding 2 (default)](https://ai.google.dev/gemini-api/docs/models/gemini-embedding-2) | General-purpose semantic search, document retrieval, and recommendations, including multimodal retrieval across text, images, video, audio, and PDFs. |
| [Gemini Embedding 001](https://ai.google.dev/gemini-api/docs/models/gemini-embedding-001) | Legacy text-only predecessor. Use it only for tables that already store its vectors. |
| [Text Embedding 3 Small](https://developers.openai.com/api/docs/models/text-embedding-3-small) | Cost-sensitive or high-volume text embedding workloads. |
| [Text Embedding 3 Large](https://developers.openai.com/api/docs/models/text-embedding-3-large) | Higher-quality text retrieval when accuracy matters more than cost. |

### Typed decision models

Typed decision models return a structured judgment, not generated prose. You give the app some text (a ticket, a comment, a list of candidates) and the questions you want answered. The model returns a category, a score, or a yes/no probability that your app can act on in code.

Use them when you need classification, routing, ranking, or verification. Keep writing replies, summaries, and free-form extraction on a [chat model](#chat-models).

Ask Lovable to:

* Route support tickets to a team or queue, with a catch-all when nothing fits.
* Flag comments against several independent policy rules.
* Rank search results, products, or candidates by relevance.
* Score quality on separate dimensions, then combine the scores in your app.
* Check a proposed answer against source text you already retrieved.

| Model | Best for |
| - | - |
| [Jev Latest (default)](https://docs.typesafe.ai/concepts/system-one) | Structured decisions from text: classification, routing, scoring, and yes/no probabilities. |

Jev evaluates text only. For images, audio, or video, transcribe or describe the content first, or use a chat or embedding model that accepts that media. It is trained primarily on English, so test other languages before you rely on them.

### Text-to-speech models

Text-to-speech models turn text into natural-sounding spoken audio, so your app can read content aloud, narrate generated text, or talk back. Describe the voice feature you want, and Lovable wires up the backend and picks the right settings. Speech streams as it is generated by default, so playback can begin before the full clip is ready, and you can describe the tone or pacing you want in plain language (for example, "speak slowly and warmly").

Ask Lovable to:

* Add a "read aloud" button that speaks an article or summary.
* Narrate AI-generated stories, lessons, or briefings.
* Turn a book or document into an audiobook.
* Build a voice assistant that responds with spoken audio.

| Model | Best for |
| - | - |
| [GPT-4o Mini TTS (default)](https://platform.openai.com/docs/models/gpt-4o-mini-tts) | Natural-sounding speech for narration, read-aloud, and voice features. |
| [Gemini 2.5 Flash TTS](https://ai.google.dev/gemini-api/docs/speech-generation) | Cost-effective speech for everyday and high-volume voice features. |
| [Gemini 2.5 Pro TTS](https://ai.google.dev/gemini-api/docs/speech-generation) | Higher-fidelity speech when voice quality matters most. |
| [Gemini 2.5 Flash Lite Preview TTS (deprecated)](https://ai.google.dev/gemini-api/docs/speech-generation) | Lightweight preview model for simple speech tasks. |
| [Gemini 3.1 Flash TTS Preview](https://ai.google.dev/gemini-api/docs/speech-generation) | Newest Gemini speech model, in preview. |

### Speech-to-text models

Speech-to-text models turn spoken audio into text, so people can talk to your app instead of typing and your app can work with what they said. Upload a voice note, call recording, or meeting audio, and the app transcribes it. Transcription streams as it is produced by default; because these features run in real time, very long recordings may time out, so transcribe long audio in segments.

Ask Lovable to:

* Build a meeting assistant that turns a recording into notes and action items.
* Add voice input or dictation so users can speak instead of type.
* Transcribe and search voice memos.
* Caption or subtitle uploaded audio.

| Model | Best for |
| - | - |
| [Gemini 3.5 Transcribe (default)](https://ai.google.dev/gemini-api/docs/models/gemini-3.5-transcribe) | Transcription with automatic language detection across more than 85 languages. Handles filler words and self-corrections. |
| [GPT-4o Mini Transcribe](https://platform.openai.com/docs/models/gpt-4o-mini-transcribe) | Fast, cost-effective transcription for most voice input and audio features. |
| [GPT-4o Transcribe](https://platform.openai.com/docs/models/gpt-4o-transcribe) | Higher-accuracy transcription when quality matters more than cost. |

Gemini 3.5 Transcribe accepts audio files up to 14 MB. For larger files, use one of the OpenAI transcription models.

For a two-way spoken conversation, such as a voice assistant or a live translator, use a [live voice model](#live-voice-models) rather than combining text-to-speech with speech-to-text.

### Live voice models

Live voice models let the people who use your app talk with it in a two-way spoken conversation. They listen and answer in real time, let people interrupt, and can have your app's backend do work during the call, such as looking up a record or saving a booking, and then speak the result. Describe the conversation you want, and Lovable sets up the voice connection and the backend work for you.

<Note>
  This section is about voice features inside the apps you build. Looking to talk to Lovable by voice while you build? See [voice conversations in Chats](/features/chats#have-a-live-voice-conversation) and [dictation in the project chat](/features/projects/chat#the-prompt-box). Those are different features.
</Note>

Ask Lovable to:

* Build a voice tutor or coach that talks with the learner and adapts to the answers it hears.
* Add a spoken assistant that looks up orders, bookings, or account details while the person is talking.
* Build a hands-free helper that guides someone through a task step by step.

| Model | Best for |
| - | - |
| [GPT-Live 1](https://platform.openai.com/docs/models/gpt-live-1) | Real-time voice conversations with interruptions and backend actions. |

Live voice works in apps that use TanStack Start, both in the preview and in your published app. For an older React + Vite app, Lovable offers to [upgrade it to TanStack Start](/features/upgrade-to-tanstack-start) first. Live voice does not need Lovable Cloud or a user login. A call ends automatically after two minutes without speech, and after one hour at most.

## Monitor AI usage and activity

Every project has an AI activity dashboard under **Cloud → AI** that shows what your app's AI features cost and how they are performing. Use it to track spend, spot failed requests, and inspect individual AI calls. Lovable reads this activity too, so the agent can help you debug failures and improve your app's AI features.

<Note>
  The **Cloud → AI** dashboard is a per-project view for monitoring and debugging AI features. It shows individual requests with their status, duration, models, token usage, cost, and, when request content is available, the redacted request and response.

  To review AI gateway spend across your workspace, go to **Workspace settings → Plans & credit usage → Usage details** and select **Run credits**. There, you can see AI gateway usage as billed credits and filter usage by project.
</Note>

Choose a time range to summarize recent AI activity. Three cards show:

* **Total cost**: the credits your app's AI requests used in the selected range.
* **Success rate**: the percentage of requests that completed successfully.
* **Avg. Duration**: the average time a request took, in milliseconds.

How far back you can view activity depends on your plan: Free workspaces can view the last 24 hours, and paid plans can view the last 90 days.

### Recent requests

The activity list shows recent AI requests, newest first. Each entry shows its status, a title taken from the request, when it ran, the model used, the input and output tokens, the credits it cost, and how long it took. When an AI action takes several steps, they appear together as a single run so you can see the cost of each step.

Lovable always records this summary information for every AI request, so both the dashboard and the agent can see each call's status, model, tokens, cost, and duration, even when the request content is not available.

### Let Lovable debug and improve your AI features

Give Lovable visibility into your app's AI calls so it can help you make them better. When this is on, Lovable keeps the full request and response for each AI call, so the agent (and you) can open a request to see exactly what was sent and returned. The agent can use this to debug failures, refine your prompts and knowledge, and reduce cost and latency. Secrets are removed automatically, and details are kept for 90 days.

It is on by default on Free and Pro. On Business and Enterprise it is off by default; enable it from the prompt on **More → AI**, or from the **AI app context** setting in [project settings](/features/projects/settings). Changing it requires permission to edit the project.

When it is off, the dashboard and the agent still see summary metrics (status, model, tokens, cost, and duration); they just cannot open a request to see its full content.

## Usage and pricing

<Note>
  **Temporary offering, subject to change:** Free, Pro, and Business workspaces receive a **4-credit monthly AI grant** for AI gateway usage in deployed apps. On Free plans, the grant resets on the 1st of each calendar month at **00:00 UTC**. On Pro and Business plans, it refreshes with the subscription billing cycle. The grant does not roll over.
</Note>

AI gateway usage is measured when AI features inside your deployed app make model calls. These requests are separate from the Lovable agent that helps you plan, build, and edit your project.

AI gateway usage uses credits. Credit usage depends on the model used and the amount of work performed, such as text tokens, generated images, seconds of generated video, finished music tracks, audio processing, call volume, or other provider-reported usage. Live voice conversations use credits for the connected time of each call, including silence and the time the model waits for your backend, with a 15-second minimum charged when the call connects. The AI activity dashboard shows this as billable voice time for each call.

AI gateway usage rates are based on the underlying provider model costs. To estimate relative model costs, refer to the official provider pricing sources linked from the [supported model list](#supported-models-for-ai-features-in-your-app). Chat requests that use [priority processing](#faster-responses-with-priority-processing) cost more than the same request served at the standard tier.

On Free, Pro, and Business plans, AI gateway usage draws from the monthly AI grant first. After that, it draws from general credits where available.

To review AI gateway usage, go to **Workspace settings → Plans & credit usage → Usage details** and select **Run credits**. For more information about AI gateway costs, monthly AI grants, top-ups, auto top-up, alerts, and usage tracking, see [Credits and usage](/introduction/credits-and-usage).

### Canceled requests

If your app cancels an in-flight AI request, some usage may still be counted. The built-in AI connector waits briefly for the provider to finish and report final usage before closing the connection.

Provider behavior on canceled requests varies, so some usage may still be billed even if the app closes the connection before the response finishes.

## Workspace rate limits

To ensure reliable performance and fair access for all users, the built-in AI connector applies rate limits per workspace. These limits help maintain system stability, prevent abuse, control costs, and provide a consistent experience for everyone. Rate limits are measured in **requests (model calls) per minute**, not tokens per minute. Video generation adds a separate limit on how many clips one project can generate at the same time: one on Free plans, and ten on paid plans. Live voice adds a limit on how many voice calls can run at the same time in a workspace: five on Free plans, and 20 on paid plans.

If your app’s requests exceed the allowed rate, the server returns a `429 Too Many Requests` status code and the request will not be processed.

If your workspace runs out of credits, the server returns a `402 Payment Required` status code. You can restore access by adding credits or enabling auto top-up in **Workspace settings → Plans & credit usage**.

For more information, see [Credits and usage](/introduction/credits-and-usage).

Rate limits are more restrictive for free users, while paid plans include higher thresholds and greater flexibility.

* **Free plan users**: upgrade anytime to increase your limits.
* **Paid plan users**: contact [Lovable Support](https://lovable.dev/support) if you need additional capacity.

## FAQ

<AccordionGroup>
  <Accordion title="Is the built-in AI connector the same thing as the Lovable agent?">
    No. The built-in AI connector adds AI features to the apps you build with Lovable. The Lovable agent is what helps you build and edit your project.

    The models listed on this page are available for AI features inside your app. They are not the models Lovable uses to write, edit, or reason about your code.
  </Accordion>

  <Accordion title="Do I need my own OpenAI, Google, or other provider API key?">
    No. Lovable automatically generates and manages the API key for each project. You do not need to create a provider account, configure billing with a model provider, or paste API keys into your app.
  </Accordion>

  <Accordion title="Where do AI calls run?">
    AI calls run through a secure backend function that Lovable creates for you, which keeps credentials and prompts server-side. The exception is [live voice](#live-voice-models): the browser streams audio directly to the voice model, while your app's server starts the call and keeps the API key.
  </Accordion>

  <Accordion title="Can I choose which model my app uses?">
    Yes. Each model type has its own default, and you can ask Lovable to use a different supported model or combination of models for a specific AI feature. See [Supported models for AI features in your app](#supported-models-for-ai-features-in-your-app).
  </Accordion>

  <Accordion title="Can my app work with voice and audio?">
    Yes. The built-in AI connector includes text-to-speech (so your app can read content aloud or respond with voice), speech-to-text (so your app can transcribe voice notes, recordings, and meetings, or take voice input), and live voice models (so people can talk with your app in a spoken conversation). Ask Lovable to add a feature such as a "read aloud" button, a voice assistant, or a meeting transcriber, and it sets up the backend for you. See [Text-to-speech](#text-to-speech-models), [Speech-to-text](#speech-to-text-models), and [Live voice](#live-voice-models).

    For higher-fidelity or character voices as a core part of your product, you can also connect [ElevenLabs](/integrations/eleven-labs).
  </Accordion>

  <Accordion title="Can my app use Anthropic (Claude) models?">
    Yes. The built-in AI connector offers Claude Fable 5.1, Opus 5.5, Opus 5, Sonnet 5, and Haiku 4.5, listed under [Chat models](#chat-models), in newer apps (created from May 13, 2026, which use TanStack Start). On Enterprise plans, a workspace admin first enables [Extended-retention models](/features/privacy-and-security-settings#extended-retention-models). If your app needs a provider that isn't offered, call that provider directly from a backend edge function using your own API key stored as a [secret](/features/secrets).
  </Accordion>

  <Accordion title="What is Jev, and when should I use it instead of a chat model?">
    [Jev Latest](#typed-decision-models) is a typed-decision model from TypeSafe. It returns a category, score, or yes/no probability instead of generated text. Use it to classify tickets, route requests, rank candidates, or verify a claim against text you already have. Use a chat model when the feature needs to write a reply, summarize, or extract new wording.
  </Accordion>

  <Accordion title="Can I use my own API key so AI usage doesn't consume my credits?">
    Not with the built-in AI connector, which always runs through Lovable and bills your workspace credits. If you want usage billed to your own provider account instead, ask Lovable to call the provider's API directly from a backend edge function, with your key stored as a [secret](/features/secrets). That usage is billed by the provider, and your app only consumes regular Cloud usage for running the function.
  </Accordion>

  <Accordion title="What happens if my workspace runs out of credits?">
    If your workspace runs out of credits, AI requests return a `402 Payment Required` status code. You can restore access by adding credits or enabling auto top-up in **Workspace settings → Plans & credit usage**.

    For more information, see [Credits and usage](/introduction/credits-and-usage).
  </Accordion>
</AccordionGroup>


## Related topics

- [FAQ](/introduction/faq.md)
- [Lovable changelog](/changelog.md)
- [Project usage and costs](/features/project-usage.md)
- [Credits and usage](/introduction/credits-and-usage.md)
- [Connect your app to Fireworks AI](/integrations/fireworks.md)
