Lyria 3.5 Clip vs Pro

Compare Lyria 3.5 Clip and Pro by duration, control, iteration speed, output workflow, and the jobs each model is built to handle.

Friday, September 4, 2026Omid Saffari
Lyria 3.5 Clip vs Pro

For Lyria 3.5 Clip vs Pro, pick Clip for 30-second assets and cheap direction-finding; pick Pro only when a song must hold together across verses, choruses, bridges, or a longer edit. The current API price is $0.04 for a Clip request and $0.08 for a full-song request, but Google's live naming is messier than the query suggests.

Which One Should You Pick?

Choose Clip when the deliverable is 30 seconds or less, or when the musical direction is still unsettled. Choose the full-song model when the arrangement itself is part of the deliverable. That means Clip for social beds, stingers, loops, and auditions; full song for a product film, podcast theme with development, campaign anthem, or any track that needs an intro, verse, chorus, bridge, and ending to feel intentional.

Decision axisLyria 3 ClipLyria 3.5 full song
Current API price$0.04 per request$0.08 per request
DurationFixed 30 secondsA couple of minutes, influenced by the prompt
Deciding capabilityCheap musical-direction auditionsPrompted duration, timestamps, and section structure
Best business jobShort ads, loops, bumpers, previewsLonger edits and complete song drafts
DealbreakerCannot become a coherent full song by itselfCosts twice as much to audition one direction

Google Lyria 3 Clip is the short-work model currently documented under lyria-3-clip-preview. Its fixed 30-second MP3 is a feature when your team needs to hear eight directions before approving one, because each rejected idea is small and cheap.

Google AI for Developers page for the Lyria 3 Clip preview model
Google Lyria 3 Clip model page

Do not buy length before you have a direction. A creative director choosing between clean electronic, warm acoustic, and orchestral tension does not need several complete songs. The decision can usually be made from the hook, palette, tempo, and vocal treatment inside a short sample.

Lyria 3.5 is the full-song model currently documented under lyria-3.5. It earns the extra four cents when continuity matters: a melody must return, energy must rise, and sections must arrive at useful moments instead of feeling like separately generated fragments.

Google AI for Developers page for the Lyria 3.5 full-song model
Lyria 3.5 full-song model page

The decision flips at 60 seconds. Below that point, Clip is the cheaper direct fit. At 60 seconds, two Clip requests and one full-song request both cost $0.08. Beyond it, one successful full-song request costs less than stacking more 30-second generations, and it preserves one musical arc.

Lyria 3.5 Clip vs Pro: What Google Lyria 3.5 Means

The live API does not currently present two endpoints named Lyria 3.5 Clip and Lyria 3.5 Pro. Verified against Google's guide, changelog, model pages, and pricing page on 4 September 2026, the short option remains Lyria 3 Clip at lyria-3-clip-preview; the new full-song option is Lyria 3.5 at lyria-3.5. Google's guide still calls the latter the "Pro model" in its explanatory text, which is why the search language makes sense even though the endpoint names do not line up cleanly.

This distinction matters in production. A model label in a planning deck can be approximate. A model ID in an API request cannot. Sending the planned but undocumented lyria-3.5-clip-preview or lyria-3.5-pro-preview strings would build a workflow around names that are absent from the current live pages.

Google's 3 September release note announces lyria-3.5 in public preview for full-length generation. The same changelog records the older Clip and Pro pair on 25 March, and the current guide now pairs the surviving Clip endpoint with the new Lyria 3.5 full-song endpoint. The useful comparison is therefore short iteration versus full-song production, not two cleanly branded 3.5 tiers.

Lyria 3.5 Pricing and the 60-Second Crossover

Clip is cheaper per request; the full-song model becomes cheaper per generated minute as soon as the required track passes one minute. Google's live pricing page, verified on 4 September 2026, lists no free API tier for either option. Lyria 3 Clip Preview costs $0.04 per song, and Lyria 3.5 Full Song costs $0.08 per song.

At 100 requests, that is $4 for Clip and $8 for full song. That comparison is useful for an experimentation budget, but it is not a fair output comparison because the results have different lengths. Normalize around the asset you need:

  • A 30-second social bed costs $0.04 with Clip or $0.08 with the full-song model. Clip wins.
  • A 60-second bed assembled from two Clip generations costs $0.08. One full-song request also costs $0.08. Price ties, but the full-song request has the structural advantage.
  • A 120-second track assembled from four Clip generations costs $0.16. One successful full-song request costs $0.08. Full song wins on both cost and continuity.

Those figures assume no retries and that one full-song request reaches the target duration. They also do not pretend that two unrelated 30-second outputs join into a song. The math shows where request cost changes; the creative constraint makes the case for full song even stronger once sections must connect.

Bar chart comparing Lyria Clip and full-song API cost at 30, 60, and 120 seconds
The request-price crossover is 60 seconds, assuming no retries.

The better budget pattern is a funnel. Audition eight musical directions with Clip for $0.32, then promote two finalists to full songs for another $0.16. Total cost is $0.48. Auditioning all eight directions as full songs costs $0.64, so the hybrid saves $0.16, or 25%, before retries.

Call this the sketch-to-score gate: cheap short outputs decide what deserves length. It is not merely a four-cent optimization. It also stops reviewers from listening through complete arrangements whose basic genre, voice, or instrumentation was wrong in the first ten seconds.

Category winner: Clip below 60 seconds; Lyria 3.5 full song above 60 seconds. The tie at exactly one minute is where coherence, not request price, should make the decision.

Iteration Speed: Clip Wins the Audition

Clip wins creative iteration because its fixed scope makes every option comparable. Google has not published a latency benchmark for the two current endpoints, so "faster" here means a shorter asset to generate, review, label, and reject, not a claimed server-speed advantage.

Consider a brand team cutting a 30-second product teaser. The team still has to choose whether the soundtrack feels precise, playful, luxurious, or urgent. A complete verse and bridge add no decision value at that stage. Eight fixed-length candidates let the reviewers compare the same window across genre, BPM, key, instrumentation, and vocal treatment.

The craft bar is whether the opening lands against the edit. Listen for a clean first beat, space for dialogue, a usable rise into the product reveal, and an ending that can cut without sounding amputated. If the track is intended to loop, also check whether the tail can meet the opening without an obvious harmonic jump. Clip is ship-ready for that bounded job after editorial and rights review.

Clip becomes mood-board-only when the work depends on development over time. A promising 30-second texture does not prove that a chorus will pay off, a second verse will vary intelligently, or a two-minute arc will stay coherent. Approving it as a full-song solution would be approving a color swatch as a finished identity system.

Category winner: Lyria 3 Clip. Its limit is exactly what makes it efficient before the musical direction is approved.

Control and Coherence: Pro Wins the Arrangement

The full-song model wins once time becomes a design material. Google's music-generation guide documents section tags such as [Intro], [Verse], [Chorus], [Bridge], and [Outro], plus timestamps that tell the model when instruments, lyrics, and energy changes should arrive. You can also specify genre, instruments, BPM, key or scale, mood, custom lyrics, and intended duration.

That control is practical for a 90-second product film. The opening can leave room for narration, the beat can widen at the feature reveal, and the final phrase can resolve under the logo. A generic prompt asks for a mood. A structured prompt gives the music a job inside the timeline.

Control is not editing. Google also documents music generation as a single-turn process: you cannot take the returned song and ask the API to soften only the second chorus in a follow-up turn. The next request is a new generation, and the same prompt can produce a different result. Build review around whole takes, not surgical revisions.

That makes prompt structure more important than prompt poetry. A useful brief states the deliverable first, then the musical grammar, then the timeline:

A 90-second instrumental product-film score at 112 BPM in D minor. Muted analog bass, dry electronic drums, and one bright arpeggiated synth. Keep 0:00-0:12 sparse for narration, add percussion at 0:12, reach the main lift at 0:42, pull back at 1:05, and end cleanly by 1:30. No vocals.

The timestamps do not guarantee a keeper. They make failure legible. A reviewer can say the lift arrived late or the narration bed became crowded, then revise one instruction for the next whole take.

Lyria Full Song Generation: When Pro Pays Off

Full song pays off when one continuous musical arc is worth more than twice as many sketches. That usually means a deliverable above 60 seconds, a recurring melodic idea, timed sections, custom lyrics, or a finish that must resolve rather than loop.

The full-song model is also the safer choice when an editor needs options inside one file. A coherent intro, low-energy verse, high-energy chorus, and clean outro create multiple usable edit points without stitching unrelated generations together. MP3 is still a limited handoff compared with stems or an editable session, but one organized track is much closer to a production asset than four disconnected clips.

Category winner: Lyria 3.5 full song. The premium buys arrangement scope and control, not merely more seconds.

Which AI Music API Workflow Fits Your Job?

Use the API as a two-stage creative review system, not a slot machine. Google itself recommends iterating with the Clip endpoint before committing to a full-length generation with Lyria 3.5. The strongest workflow keeps the same musical brief across both stages and changes only the duration and structure layer.

Here is the before-to-after on one product-launch brief. The weak starting prompt is: "Make upbeat modern music for a product video." It leaves genre, tempo, tonal center, instrumentation, vocal policy, edit points, and ending unresolved. Any result could be defensible, which makes feedback vague.

The improved Clip prompt is: "A 30-second instrumental product teaser at 112 BPM in D minor, with muted analog bass, dry electronic drums, and one bright arpeggiated synth. Precise and confident, not aggressive. Start clean, lift at 12 seconds, and end on a resolved hit. No vocals." This does not claim a better result. It creates a better test because reviewers can judge a specific direction against explicit criteria.

  1. Define the acceptance test

    Write four checks before generating: the first beat supports the edit, dialogue has space, the lift lands at the reveal, and the ending cuts cleanly. Add brand-specific exclusions such as no vocals or no retro drum sounds.

  2. Audition with Clip

    Keep duration at 30 seconds. Change one major variable per round, such as genre, instrument family, or energy, while holding the product-film job constant. Label each output by the choice it tests, not by an arbitrary take number.

  3. Promote the winning direction

    Move the approved genre, BPM, key, instrument palette, and mood into the full-song prompt. Add section tags or timestamped events only now, because structure is expensive to review before the direction is approved.

  4. Inspect the handoff

    Save the MP3 and returned lyric or structure text together. Inspect the actual sample rate, loudness, start and end, safety-sensitive content, and whether every requested section is present before the asset reaches an editor.

  5. Finish outside the generator

    Trim, level, sequence, and test the track against the final picture. If the work needs stems, note that the documented API response is an MP3 plus text, not a multitrack session.

Decision flow for choosing Lyria Clip first or moving to the full-song model
Approve the direction in Clip, then buy full-song structure only when the brief needs it.

For a 30-second ad, stop after Clip if the output clears the acceptance test. For a 90-second product film, use Clip only to choose the sound, then generate the final arc with Lyria 3.5. For a very short sting, Clip is still the closest fit, but expect to edit the 30-second output down. For real-time interactive music, neither is the right branch; Google points developers to Lyria RealTime instead.

The Lyria 3.5 API Output Workflow and Its Production Gaps

The current handoff is good enough for a finished stereo asset and too thin for detailed music production. Both entries use Google's Interactions API, accept text and image inputs, and return MP3 audio plus text for lyrics or structure. They do not expose the kind of multitrack session a producer would use to rebalance vocals, drums, and bass independently.

The first production gap is editing. Each generation is single-turn, so a selected Clip cannot be extended conversationally into the same song, and a full-song take cannot be locally repaired through follow-up prompting. Treat prompts as reproducible briefs, not editable project files.

The second gap is scaled operations. The current Clip and full-song model pages list a 131,072-token input limit, but they also say Batch API, Flex inference, and Priority inference are unsupported. A product that needs a scheduled bulk catalog or guaranteed priority path should not assume the standard preview endpoints provide it.

The third gap is documentation consistency. The combined guide says the current pair produces 44.1 kHz stereo, while the still-live Clip model page says that endpoint produces 48 kHz. Do not hard-code an ingest rule from either sentence. Read the returned file's metadata, normalize it in post, and keep the original.

Google applies SynthID to Lyria audio. The company says the watermark is inaudible and designed to remain detectable through common changes such as MP3 compression, added noise, and speed adjustments. That is useful provenance, but it is not a substitute for your own asset log, prompt record, source-rights check, and disclosure policy.

Google's API terms say the company does not claim ownership of generated content, while making the user responsible for lawful use and possible attribution obligations. Paid Gemini API prompts and responses are not used to improve Google's products, although limited safety and security logging remains. For business work, keep the project on a billed Cloud project and do not mistake "Google does not claim ownership" for a promise of exclusivity or infringement clearance.

Category winner: neither. The shared Interactions API and MP3 handoff are straightforward, but both inherit preview constraints. Clip is ship-ready for bounded short assets after review. The full-song model is ship-ready for stereo drafts and lower-stakes final beds, but signature music, releases that need stems, and tightly mixed client masters still need a production layer outside the API.

If your real requirement is a reusable sonic identity rather than one-off tracks, compare this workflow with the ElevenLabs Music Finetunes API approach. If the generated track is headed into a complete visual edit, the AI music video generator comparison covers the downstream assembly problem.

Switching From Clip to Pro: What Changes

Switching models costs almost nothing in stored data and more than it appears in workflow. Both routes sit behind the Interactions API, so the basic request and response pattern stays familiar. The prompt, asset contract, and review standard change.

First, preserve the approved Clip prompt as the source of truth. Carry genre, BPM, key, instrumentation, mood, and vocal policy into the full-song brief. Then add the structure the longer asset needs: section tags, timing, energy changes, and an explicit ending.

Second, change the model ID through configuration. The short endpoint is lyria-3-clip-preview; the current full-song endpoint is lyria-3.5. Both are preview-stage names, so a hard-coded identifier across multiple services creates avoidable lock-in when Google changes the contract again.

Third, revise quality assurance. A Clip review asks whether one idea works for 30 seconds. A full-song review also asks whether motifs return coherently, transitions earn their place, lyrics remain usable, the middle avoids repetition, and the ending resolves. That extra listening and editorial judgment is the real switching cost.

Do not expect continuity from the selected Clip. Because multi-turn editing is unsupported, the full-song call is a new performance of the brief, not an extension of the exact audio. The approved Clip is a direction reference. If matching its precise melody or vocal performance is essential, this workflow does not promise that bridge.

Who should not switch? Stay with Clip when every deliverable is a short social bed, loop, bumper, or prototype. Do not build on either endpoint when you need deterministic regeneration, native stems, a documented batch lane, priority inference, or local repair of one section. Those are workflow requirements, not reasons to hope a longer model will behave like a digital audio workstation.

Benchmarks: What Is and Is Not Proven

No current benchmark proves that the full-song model sounds better than Clip on the same 30-second brief. Google's Lyria 3.5 model card says its own human and automated evaluations found significant audio-fidelity improvement over Lyria 2 and better prompt adherence for lyrics, but it publishes no numeric scores and no short-versus-full-song head-to-head.

That is a vendor evaluation, not an independent comparison. A 2026 academic audit of AI music homogenization examines Lyria 3 against Suno across several genres, not the current short and full-song pairing. Its findings should not be transplanted into this decision.

The honest buying rubric is operational: duration fit, structural adherence, useful edit points, rejection rate, and post-production burden on your own licensed brief. Until an independent evaluator publishes matched prompts and complete outputs for both current endpoints, any universal audio-quality winner would be guesswork.

The Monday Move

On Monday, choose one live 30-second campaign cut and write four acceptance criteria before opening the API. Generate eight Clip directions, changing one meaningful variable at a time. Keep the best two only if they pass the same review rubric.

If the final deliverable is short, finish with the winning Clip and move it into the edit. If the campaign also needs a longer product film or theme, promote those two briefs to the full-song model with timestamped structure. The base generation spend is $0.48 for eight auditions and two full-song finalists.

Save each MP3 beside its prompt, returned structure text, model ID, and generation date. Inspect the file metadata instead of assuming a sample rate. Judge the music against picture, not in isolation. That gives the team a repeatable decision trail even though the model itself does not offer multi-turn editing or deterministic reruns.

Frequently Asked Questions

Should You Use Lyria 3.5 Clip or Pro for Business Music?

Use Clip to approve a 30-second musical direction, short ad bed, loop, bumper, or preview. Use the current Lyria 3.5 full-song endpoint when the deliverable needs more than 60 seconds, timed sections, or one coherent arrangement.

Can Lyria Clip Generate Full Songs?

No. Google's current guide says the Clip endpoint always generates 30 seconds. Multiple clips can be edited together, but they are separate generations and do not provide the continuity of one full-song request.

How Long Can Lyria 3.5 Songs Be?

The Gemini API guide promises songs lasting a couple of minutes with duration influenced by the prompt. Google DeepMind's broader product page shows tracks up to three minutes, but that should not be treated as an exact API guarantee until the API documentation says so.

What Is the Price Difference Between Lyria Clip and Pro?

The live paid API prices are $0.04 for a Lyria 3 Clip request and $0.08 for a Lyria 3.5 full-song request, with no free API tier listed. Clip is half the request price; full song becomes cheaper per generated minute above the 60-second crossover when one request reaches the target.

Before you add AI music to a production workflow, get the AI Business Workflow Audit Checklist.

Last Updated

Sep 4, 2026

CategoryDesign

Prefer this site in Google

Add omidsaffari.com as a preferred source in Google Search

Mark omidsaffari.com as preferred and Google lifts it in Top Stories, AI Overviews and AI Mode for you.

More from Design

View all Design articles
Newsletter

One letter, every Sunday. Working systems, not hot takes.

Build logs, working systems, and field notes from running a portfolio of AI ventures.

Weekly. No spam. Unsubscribe anytime.