Apple Patents a Siri That Reads Your Photos and Texts Together
Apple wants Siri to feel less like a voice prompt and more like a text conversation, and this patent adds a twist: Siri would understand a photo you drop into the chat and then connect it to whatever you type next.
How Apple's messaging-based Siri would actually work
Right now, when you ask Siri something, each request is basically a fresh start. Send a photo to Siri today and then type a follow-up question, and it won't automatically know the two things are related.
This patent describes a version of Siri built around a messaging thread, the same back-and-forth format you use in iMessage. You could drop a photo into the conversation, type something like "add this to my calendar," and Siri would understand that your words refer to the event details in the image. The conversation history stays visible, just like a chat with a friend.
The key difference is that Siri would hold onto context between your messages, so a photo you shared two bubbles ago still informs what you're asking right now. That changes how natural the whole experience feels.
… cause a user intent to be determined based on a combination of the media object and the first text; after the user intent is determined, perform a task in accordance with the user intent; and display, in the GUI, a first response based on the task.
Translation: The system analyzes your photo and your text message together to figure out what you want and then completes the request.
How Siri links a photo to your follow-up text
The patent describes a system where the digital assistant operates inside a graphical user interface (GUI) formatted exactly like a standard messaging app, with conversation bubbles for both the user and the assistant.
Here's how a typical interaction would work:
- A user shares a media object, such as a photo or image, directly into the chat window.
- The user then types a text message, such as a question or command, as a follow-up.
- The system determines a user intent by combining both inputs, the image and the text, rather than treating them as separate, unrelated requests.
- Siri performs the appropriate task and replies with a message in the same conversation thread.
Critically, the patent also describes storing a contextual state at the moment each message is sent. That means the system can remember what was on the screen, what media was shared, and what had been said earlier in the conversation, preserving that context so follow-up messages stay coherent.
This multi-modal input approach (handling both images and text together) is the core engineering claim. The assistant doesn't just respond to words; it responds to the full picture of what the user shared.
Systems and processes for operating an intelligent automated assistant in a messaging environment are provided. In one example process, a graphical user interface (GUI) having a plurality of previous messages between a user of the electronic device and the digital assistant can be displayed on a display.
Translation: This patent describes how a digital assistant can hold a conversation with you inside a standard messaging app.
What this means for how you talk to Siri every day
For most people, the friction with voice or text assistants isn't any single broken feature; it's the constant need to re-explain yourself. You have to front-load every request with full context because the assistant forgets everything the moment a task ends. A persistent, context-aware thread would let you build on earlier messages the same way you do in a text conversation with another person.
Apple has been steadily reworking how Siri handles multi-step and context-dependent requests, and this filing fits that broader push. For anyone tracking Big Tech patent news around AI assistants and on-device intelligence, this patent maps how Apple is thinking about blending the messaging habit people already have with the assistant experience they want Siri to become.
This is the sixth Apple filing we've tracked since June in our AI assistants that remember you watch, following one on reading your app activity and one on summoning AI in virtual worlds.
You know the moment: you sent a photo of a menu item to your phone's assistant, asked a follow-up question thirty seconds later, and it had no idea what you were talking about. You had to start over, re-explain everything, and by then the moment had passed.
Apple's patent describes an assistant that lives inside a chat thread and holds onto everything you've shared across the whole conversation, the photos, the questions, the context, treating each new message as a continuation rather than a fresh start.
For most people that means fewer repeated explanations, faster answers, and an assistant that actually earns its place in the day instead of adding friction to it.
There are more where this came from
We read every patent application Big Tech publishes and send you the ones worth knowing. Plain English, free, every week.
The drawings
41 drawing sheets from US 2026/0252370 A1 · click any drawing to enlarge
Want this weekly breakdown for a company we don't cover? Patentlyze Pro →