← All release notes
SwiftiOS 26Aug 17, 2026 · 5 min read

The Foundation Models Framework — On-Device LLM Access for Apps

iOS 26's Foundation Models framework gives apps on-device LLM access with typed, structured output — no API key, no network call.

iOS 26 opens up the on-device large language model behind Apple Intelligence through the Foundation Models framework.

It's a real model with real limits. There are no server calls, it isn't an open-ended chat product, and it's far smaller than the models behind server-scale assistants. For structured, in-app tasks, though, it deletes a pile of boilerplate. No API key to manage, no network dependency, and nothing billed per token.

#Checking availability first

The model runs only on Apple Intelligence–eligible devices, with the feature switched on, in a supported region. Every integration opens with an availability check instead of assuming the model is there.

import FoundationModels

let model = SystemLanguageModel.default

switch model.availability {
case .available:
    break
case .unavailable(let reason):
    print("Model unavailable: \(reason)")
}

Branch your UI on this rather than force-using the model. Devices without Apple Intelligence, or with it turned off, are not an edge case you can ignore.

#Sessions hold the conversation

A LanguageModelSession is the object you talk to. It keeps prompt and response history across calls, like a chat thread. You can also hand it optional instructions at creation, which act as a system prompt.

let session = LanguageModelSession {
    "You are a concise assistant for a travel app. Keep answers under two sentences."
}

let response = try await session.respond(to: "Suggest a 3-day itinerary theme for Lisbon.")
print(response.content)

One session per conversation or task, not one per request. Reusing a session across turns is exactly what gives you multi-turn context.

#Guided generation: typed output instead of parsed strings

Guided generation is the standout. Mark a Swift type @Generable and the model hands back that type directly, populated and schema-validated, instead of a string you regex apart. @Guide annotations narrow individual fields further.

@Generable
struct TripSuggestion {
    let destination: String

    @Guide(description: "A short, upbeat tagline under ten words")
    let tagline: String

    @Guide(.anyOf(["budget", "moderate", "luxury"]))
    let budgetTier: String
}

let suggestion = try await session.respond(
    to: "Suggest a European city for a first-time solo traveler.",
    generating: TripSuggestion.self
)
print(suggestion.content.destination)

Use it anywhere you'd otherwise pick apart free text: categorization, pulling structured fields out of a description, filling a form. Beats prompting for JSON and decoding it yourself.

#Streaming partial results

Longer generations get a streaming variant. It yields partial values as the model produces them instead of making the caller wait for the whole response.

let stream = session.streamResponse(generating: TripSuggestion.self) {
    "Suggest a European city for a first-time solo traveler."
}

for try await partial in stream {
    updateUI(with: partial)
}

Reach for this when a blank screen followed by a sudden full response would feel broken. Multi-paragraph text especially.

#Tool calling

A session can call into your own app code mid-generation through the Tool protocol. Declare a name, a description, a @Generable arguments type, and an async call(arguments:) method returning a ToolOutput. Pass tools in at session creation, and the model decides when to invoke them based on the prompt and each tool's description.

That's how you ground the model in live data it never saw in training, instead of stuffing everything into the prompt text. A tool can run a local database lookup, report the current date, or read in-app state.

#Handling errors and guardrails

This framework has its own failure modes on top of the usual network errors. A request can trip the built-in content guardrails, blow past the model's context window, or hit an unsupported language. Wrap calls in do/catch and branch on the specific failure rather than showing one generic error for all of them.

do {
    let response = try await session.respond(to: userPrompt)
    print(response.content)
} catch {
    print("Generation failed: \(error)")
}

The framework's error type distinguishes these cases. Worth inspecting rather than swallowing, since a guardrail rejection and a context-window overflow call for different UI responses. Exact case names have shifted across betas, so check the current definition rather than trusting an older blog post's list verbatim.

#What to adopt first

  • Ship the availability check everywhere before anything else. It isn't optional plumbing. On devices without Apple Intelligence it is the majority of your integration work.
  • Start with guided generation on a narrow, low-stakes task before reaching for open-ended chat. Categorize or summarize existing in-app text, generate short copy. That plays to what the on-device model is actually good at.
  • Don't design a feature that assumes every user has this model. Apple Intelligence availability, plus devices below the deployment target, is a meaningful chunk of any real install base.
  • Keep the model honest about app-specific facts with tool calling rather than prompt engineering. For anything the model can't know from training data, give it a tool rather than hoping it doesn't hallucinate.