2:26-cv-00372
Cerence Operating Co v. Amazon.com Inc
I. Executive Summary and Procedural Information
- Parties & Counsel:
- Plaintiff: Cerence Operating Company (Delaware)
- Defendant: Amazon.com, Inc., Amazon.com Services LLC, and Amazon Web Services, Inc. (Delaware)
- Plaintiff's Counsel: TROUTMAN PEPPER LOCKE LLP
- Case Identification: 2:26-cv-00372, E.D. Tex., 05/04/2026
- Venue Allegations: Venue is alleged based on Defendants having regular and established places of business in the Eastern District of Texas, including numerous fulfillment centers and delivery stations, and deriving substantial revenue from sales of the Accused Products within the district.
- Core Dispute: Plaintiff alleges that Defendant's Amazon Alexa virtual assistant, and the devices and cloud services that incorporate it, infringe five patents related to conversational AI, speech recognition, and audio processing technologies.
- Technical Context: The lawsuit concerns core technologies in the rapidly expanding market for voice-activated virtual assistants and smart devices, focusing on user experience and system efficiency.
- Key Procedural History: Plaintiff Cerence Operating Company was spun out from Nuance Communications, Inc., a long-standing leader in speech and language solutions, in October 2019. This history may be relevant to establishing the patents' lineage and the company's long-standing innovation in the asserted technical fields.
Case Timeline
| Date | Event |
|---|---|
| 2006-12-14 | '815 Patent Priority Date |
| 2007-01-08 | '484 Patent Priority Date |
| 2007-10-01 | '575 and '972 Patents Priority Date |
| 2012-11-06 | U.S. Patent No. 8,306,815 Issues |
| 2012-11-27 | U.S. Patent No. 8,320,575 Issues |
| 2013-01-15 | U.S. Patent No. 8,355,484 Issues |
| 2015-12-01 | U.S. Patent No. 9,203,972 Issues |
| 2019-03-28 | '073 Patent Priority Date |
| 2019-10-01 | Cerence becomes a separate public company after spin-off from Nuance |
| 2024-03-12 | U.S. Patent No. 11,929,073 Issues |
| 2026-05-04 | Complaint Filed |
II. Technology and Patent(s)-in-Suit Analysis
U.S. Patent No. 8,355,484 - Methods and Apparatus for Masking Latency in Text-to-Speech Systems, Issued Jan. 15, 2013
The Invention Explained
- Problem Addressed: The patent describes that in an automatic dialog system, the compounded processing latencies of speech recognition, natural language understanding, and response generation can lead to frustrating delays between a user's query and the system's spoken reply '484 Patent, col. 1:26-34
- The Patented Solution: The invention proposes masking this latency by having the system provide at least one "transitional message" while the main response is being processed '484 Patent, abstract These transitional messages can be paralinguistic events (e.g., "uhm," "hmmm") or canned phrases (e.g., "let me see"), which mimic natural human conversational pauses and inform the user that the system is working, thereby making the delay feel more natural and less like a system failure '484 Patent, col. 2:13-20 '484 Patent, FIG. 2
- Technical Importance: This approach aims to improve the user experience of conversational AI by managing user perception of system latency, a critical factor for user acceptance of dialog systems.
Key Claims at a Glance
- The complaint asserts infringement of the patent's "independent asserted claims" without specifying claim numbers Compl. ¶31 Assuming assertion of the first independent claim, Claim 1 is a method claim.
- Key Elements of Claim 1:
- Receiving a communication from a user at an automatic dialog system.
- Processing the communication to provide a response.
- Providing at least one "transitional message" to the user while processing the communication.
- The transitional message comprises at least one of a paralinguistic event and a phrase.
- Providing the response to the user after providing the transitional message.
- The complaint does not explicitly reserve the right to assert dependent claims for this patent.
U.S. Patent No. 8,306,815 - Speech Dialog Control Based on Signal Pre-Processing, Issued Nov. 6, 2012
The Invention Explained
- Problem Addressed: The patent notes that the reliability of speech dialog systems can degrade significantly in noisy environments, such as inside a vehicle, which can frustrate the user and render the system ineffective '815 Patent, col. 1:26-33
- The Patented Solution: The invention describes a system with a "signal pre-processor" that generates two separate outputs from a user's speech input: an "enhanced speech signal" (for recognition) and an "analysis signal" '815 Patent, abstract This analysis signal contains non-semantic information about the speech input, such as background noise level, echo, speaker location, or volume '815 Patent, col. 2:34-39 A "control unit" uses this analysis signal to intelligently manage the dialog-for example, by increasing the system's output volume in a noisy environment or asking the user to speak louder '815 Patent, col. 6:11-20 '815 Patent, FIG. 4
- Technical Importance: This technology allows a speech system to adapt its behavior in real-time to the acoustic environment, improving robustness and usability beyond simple noise filtering.
Key Claims at a Glance
- The complaint asserts infringement of the patent's "independent asserted claims" without specifying claim numbers Compl. ¶38 Assuming assertion of the first independent claim, Claim 1 is a system claim.
- Key Elements of Claim 1:
- A signal pre-processor unit to process a speech input signal and generate both an enhanced speech signal and an analysis signal, where the analysis signal contains non-semantic characteristics.
- A speech recognition unit to receive the enhanced signal and generate a recognition result.
- A speech output unit.
- A speech dialog control unit that receives the analysis signal and the recognition result, and controls the speech output unit based on the analysis signal.
- The complaint does not explicitly reserve the right to assert dependent claims for this patent.
U.S. Patent No. 8,320,575 - Efficient Audio Signal Processing in the Sub-Band Regime, Issued Nov. 27, 2012
Technology Synopsis
The patent addresses the high computational demand of audio signal processing (e.g., for echo cancellation) '575 Patent, col. 1:26-33 The solution involves dividing an audio signal into multiple frequency "sub-bands," "excising" (deleting) a subset of these bands to reduce the amount of data, processing the remaining bands, and then reconstructing the excised bands from the information in the remaining bands before synthesizing the final enhanced audio signal '575 Patent, abstract '575 Patent, col. 2:7-14 This method aims to reduce processing load while maintaining acceptable audio quality.
Asserted Claims
The complaint asserts infringement of the patent's "independent asserted claims" without specifying numbers Compl. ¶45
Accused Features
The complaint accuses Amazon's Echo devices, smart displays, and other Alexa-enabled products of infringing Compl. ¶44 The alleged infringement presumably relates to the audio processing pipelines used for features like echo cancellation in these devices.
U.S. Patent No. 9,203,972 - Efficient Audio Signal Processing in the Sub-Band Regime, Issued Dec. 1, 2015
Technology Synopsis
This patent is a divisional of the application that led to the '575 Patent and shares the same specification '972 Patent, Related U.S. Application Data It describes the same technology for efficiently processing audio signals by excising and later reconstructing frequency sub-bands to reduce computational load '972 Patent, abstract
Asserted Claims
The complaint asserts infringement of the patent's "independent asserted claims" without specifying numbers Compl. ¶52
Accused Features
The complaint accuses the same range of Amazon Alexa products, presumably for using the patented sub-band processing techniques for audio enhancement and echo cancellation Compl. ¶51
U.S. Patent No. 11,929,073 - Hybrid Arbitration System, Issued Mar. 12, 2024
Technology Synopsis
The patent addresses challenges in hybrid (on-device and cloud) speech recognition systems, where the two processing paths often produce different results '073 Patent, col. 1:40-45 The invention is a method for an "arbitrator" to decide whether to use the fast result from the on-device ("embedded") recognizer or wait for the potentially more accurate but slower result from the cloud service '073 Patent, abstract '073 Patent, col. 2:7-12 The decision is made by a classifier that uses not just confidence scores but also raw ASR strings and NLU features from the on-device result to determine if it is "good enough" to use immediately '073 Patent, col. 2:65-col. 3:2
Asserted Claims
The complaint asserts infringement of the patent's "independent asserted claims" without specifying numbers Compl. ¶59
Accused Features
The complaint accuses Amazon's Echo devices and other Alexa products of infringement, suggesting their hybrid on-device and cloud-based architecture for Alexa voice processing utilizes the claimed arbitration method Compl. ¶58
III. The Accused Instrumentality
Product Identification
The "Accused Products" are broadly defined to include Amazon-branded smart devices with Amazon Alexa, such as Echo smart speakers and displays, Fire TVs, Fire Tablets, and the Amazon Echo Auto, as well as the underlying "Alexa cloud services" that power them Compl. ¶1
Functionality and Market Context
The complaint alleges that these products provide a voice-activated virtual assistant experience Compl. ¶26 Functionally, they receive spoken commands from a user, process the commands using speech recognition and natural language understanding, and provide a response. The complaint notes the expansion of this technology from its automotive origins into home entertainment and consumer electronics, where it allows for control over content, smart home devices, and information retrieval Compl. ¶26 The ubiquity of the Accused Products positions them as a major force in the consumer virtual assistant market.
IV. Analysis of Infringement Allegations
The complaint references claim chart exhibits for each asserted patent Compl. ¶31 Compl. ¶38 Compl. ¶45 Compl. ¶52 Compl. ¶59, but these exhibits were not provided with the complaint. The infringement theories are therefore summarized based on the complaint's narrative allegations. No probative visual evidence provided in complaint.
'484 Patent Infringement Allegations
The complaint alleges that Amazon's products infringe by providing transitional feedback while processing a user's request, which serves to mask system latency Compl. ¶¶29-30 This suggests that when an Alexa-enabled device shows a light-ring animation or emits a sound (e.g., a "hmmm" or a short tone) after a user speaks but before delivering the final answer, Cerence will argue this functionality constitutes the "transitional message" required by the claims.
- Identified Points of Contention:
- Scope Questions: A central question will be whether the visual and/or audio cues used by Alexa devices qualify as a "transitional message" comprising a "paralinguistic event" or "phrase" as defined by the patent. The defense may argue that simple processing indicators, like a spinning light, are not "messages" or "phrases" within the claim's meaning.
- Technical Questions: Evidence will be required to show that the purpose of Alexa's cues is to mask latency, as opposed to simply indicating that the device has heard the user and is processing the request.
'815 Patent Infringement Allegations
The complaint alleges that Amazon's products infringe by using signal pre-processing to control the speech dialog Compl. ¶¶36-37 The theory appears to be that Alexa-enabled devices analyze ambient conditions (e.g., background noise) and use that information to alter the system's behavior (e.g., by adjusting the volume of the spoken response). This functionality is alleged to map to the patent's system of using an "analysis signal" to inform a "control unit."
- Identified Points of Contention:
- Scope Questions: The dispute may focus on whether Amazon's system architecture aligns with the specific structure of the claims, which require a "signal pre-processor" that generates a distinct "analysis signal" used by a "speech dialog control unit."
- Technical Questions: A key evidentiary issue will be demonstrating that the accused devices generate and use a data signal corresponding to the claimed "analysis signal," containing non-semantic information, as opposed to achieving a similar adaptive result through a different, non-infringing architecture.
V. Key Claim Terms for Construction
For the '484 Patent
- The Term: "transitional message"
- Context and Importance: This term is the core of the invention for masking latency. Its construction will determine whether the accused system's sounds, lights, or brief utterances constitute infringement. Practitioners may focus on this term because it sits at the boundary between a functional processing indicator and a communicative, human-like utterance.
- Intrinsic Evidence for Interpretation:
- Evidence for a Broader Interpretation: The claim itself defines the term broadly as comprising "at least one of a paralinguistic event and a phrase" '484 Patent, claim 1 The specification provides examples such as "paralinguistic events (e.g., 'uh', 'um', 'hmmm') or canned phrases (e.g., 'let me see')" '484 Patent, col. 2:17-18, which could support reading the term to cover a wide range of non-substantive system outputs.
- Evidence for a Narrower Interpretation: The background section states a goal is for the dialog system to "act similar to a human speaker" '484 Patent, col. 1:46-47 A defendant may argue this context requires the "transitional message" to be more than a simple electronic tone or light, but rather something that genuinely mimics human conversational fillers.
For the '815 Patent
- The Term: "analysis signal"
- Context and Importance: This term is crucial for distinguishing the invention from general-purpose audio enhancement. Infringement hinges on whether the accused system generates a discrete signal with the specific characteristics and purpose recited in the claims. The debate will likely center on whether Amazon's architecture has an equivalent component.
- Intrinsic Evidence for Interpretation:
- Evidence for a Broader Interpretation: The specification provides a non-exhaustive list of information the signal may contain, using "and/or" to connect the elements, suggesting flexibility: "information related to one or more of the following non-semantic characteristics: (1) a noise component ...; (2) an echo component ...; (3) a location of a source ...; (4) a volume level ...; (5) a pitch ...; and/or (6) a stationarity" '815 Patent, col. 2:34-39
- Evidence for a Narrower Interpretation: The patent's block diagram (FIG. 2) depicts the "analysis signal" as a distinct output path from the "Signal Pre-processing Unit" (202) to the "Speech Dialog Control Unit" (206), separate from the path of the "enhanced speech signal" to the "Speech Recognition Unit" (204). This figure may support an argument that the claims require a structurally separate signal, not merely metadata embedded within a single audio stream.
VI. Other Allegations
- Indirect Infringement: For all asserted patents, the complaint alleges induced infringement, stating that Amazon provides products "with knowledge and specific intent to cause that infringement" Compl. ¶32 Compl. ¶39 Compl. ¶46 Compl. ¶53 Compl. ¶60 This is supported by allegations that Amazon disseminates product information and online guides that instruct end users on how to use the infringing features Compl. ¶18 Compl. ¶9, fn. 3 The complaint also alleges contributory infringement, stating the products are not staple articles of commerce suitable for substantial noninfringing use Compl. ¶33 Compl. ¶40 Compl. ¶47 Compl. ¶54 Compl. ¶61
- Willful Infringement: The complaint alleges that Amazon's infringement "is and has been willful" Compl. ¶34 Compl. ¶41 Compl. ¶48 Compl. ¶55 Compl. ¶62 The factual basis appears to rest on knowledge acquired at least as of the filing date of the complaint, which would support a claim for post-suit willfulness Compl. ¶32 Compl. ¶39 Compl. ¶46 Compl. ¶53 Compl. ¶60
VII. Analyst's Conclusion: Key Questions for the Case
- A question of definitional scope: Can Cerence prove that the various sounds, lights, and brief on-screen texts used by Amazon's Alexa platform during processing constitute a "transitional message" as claimed in the '484 patent, or will the court construe the term more narrowly to require specific, human-like conversational fillers?
- A question of architectural mapping: Does Amazon's hybrid on-device/cloud architecture for Alexa align with the specific system structures recited in the '815 and '073 patents? The case may turn on whether Amazon's implementation of adaptive audio and hybrid processing can be shown to contain discrete components that function as the claimed "analysis signal" and "arbitrator," respectively, or if Amazon achieves similar outcomes through a fundamentally different and non-infringing design.
- A question of technical implementation: For the '575 and '972 patents, a key evidentiary issue will be whether Amazon's audio processing for features like echo cancellation relies on the claimed method of excising and reconstructing frequency sub-bands to improve efficiency. Cerence will need to present evidence from reverse engineering or discovery showing this specific technique is practiced, while Amazon may argue it uses alternative, more common methods for computational efficiency.