1:23-cv-00581
Dialect LLC v. Amazon.com Inc
I. Executive Summary and Procedural Information
- Parties & Counsel:
- Plaintiff: Dialect, LLC (Texas)
- Defendant: AMAZON.COM, INC., and Amazon Web Services, Inc. (Delaware)
- Plaintiff’s Counsel: Hausfeld LLP; BLUE PEAK LAW GROUP LLP
- Case Identification: Dialect, LLC v. AMAZON.COM, INC., 1:23-cv-00581, E.D. Va., 07/31/2023
- Venue Allegations: Plaintiff alleges venue is proper in the Eastern District of Virginia because Amazon maintains a regular and established place of business in the district, including its second corporate headquarters (HQ2) in Arlington, Virginia, where teams working on the accused Alexa technology are located.
- Core Dispute: Plaintiff alleges that Defendant’s Alexa-enabled products and services infringe seven patents related to natural language understanding, conversational voice recognition, and multi-modal user interfaces.
- Technical Context: The technology at issue involves foundational principles of natural language understanding (NLU) and conversational AI, which enable devices to interpret and respond to human speech in a flexible, context-aware manner.
- Key Procedural History: The complaint alleges that Defendant Amazon had pre-suit knowledge of the asserted patent portfolio through a series of business and technical presentations given by the original inventor, VoiceBox Technologies, beginning in 2011. The complaint further alleges that Amazon subsequently hired VoiceBox’s Chief Scientist and recruited dozens of its engineers and scientists before and after launching the accused Alexa products.
Case Timeline
| Date | Event |
|---|---|
| 2001-01-01 | VoiceBox Technologies founded by the Kennewick brothers. |
| 2002-06-03 | Earliest Priority Date for ’006 and ’327 Patents. |
| 2002-07-15 | Earliest Priority Date for ’720, ’468, ’845, ’039, and ’957 Patents. |
| 2010-04-06 | U.S. Patent No. 7,693,720 issues. |
| 2011-09-06 | U.S. Patent No. 8,015,006 issues. |
| 2011-10-07 | VoiceBox teleconference with Amazon corporate development. |
| 2011-10-19 | VoiceBox meeting with Amazon personnel at Amazon's offices. |
| 2011-10-26 | Amazon personnel meeting at VoiceBox's office. |
| 2012-03-20 | U.S. Patent No. 8,140,327 issues. |
| 2012-06-05 | U.S. Patent No. 8,195,468 issues. |
| 2013-01-01 | IEEE ranks VoiceBox #13 in patent power for computer software. |
| 2014-01-01 | Amazon announces the launch of Alexa and the first-generation Echo. |
| 2015-05-12 | U.S. Patent No. 9,031,845 issues. |
| 2016-01-01 | Amazon hires VoiceBox's Chief Scientist, Philippe Di Cristo. |
| 2016-02-16 | U.S. Patent No. 9,263,039 issues. |
| 2016-11-15 | U.S. Patent No. 9,495,957 issues. |
| 2017-01-10 | Amazon hosts an invite-only networking event for VoiceBox employees. |
| 2017-01-17 | VoiceBox CEO sends a letter to Amazon CEO Jeff Bezos regarding employee poaching. |
| 2023-07-31 | Amended Complaint filed. |
II. Technology and Patent(s)-in-Suit Analysis
U.S. Patent No. 7,693,720 - “Mobile Systems And Methods For Responding To Natural Language Speech Utterance”
- Issued: April 6, 2010 (Compl. ¶40).
The Invention Explained
- Problem Addressed: The patent’s background section describes the difficulty of creating natural language speech interfaces for vehicles (Compl. ¶43). Conventional systems required "highly structured" commands that were not natural for users and could not adequately handle the ambiguity and incompleteness inherent in human speech (Compl. ¶¶42-44; ’720 Patent, col. 1:34-58).
- The Patented Solution: The invention is a mobile system designed to overcome these limitations through a novel architecture (Compl. ¶45). It employs a speech recognition engine that uses dictionaries and phrase lists that are "dynamically updated based on at least a history of a current dialog" to better understand the user (Compl. ¶45; ’720 Patent, col. 32:7-11). It also uses a system of "domain agents" to determine the context of an utterance, select the appropriate agent for the task, and formulate a command in a grammar that the agent can process (’720 Patent, col. 4:5-14; Compl. ¶45). Figure 5 of the patent illustrates this architecture, showing a speech unit feeding a speech recognition engine and parser, which in turn interact with various agents and databases (Compl. ¶46; ’720 Patent, fig. 5).
- Technical Importance: The technology represented a step away from rigid "command and control" voice systems toward a more fluid, context-aware conversational interface suitable for the demanding mobile environment (Compl. ¶¶43-44).
Key Claims at a Glance
- The complaint asserts independent Claim 1 (Compl. ¶104).
- The essential elements of Claim 1 include:
- A mobile system responsive to a user’s natural language speech utterance.
- A speech unit on a vehicle that receives the speech and converts it to an electronic signal.
- A natural language speech processing system that processes the signal using data from a plurality of domain agents.
- The system includes a speech recognition engine that uses dictionary and phrase entries which are dynamically updated based on a history of a current dialog and prior dialogs.
- The system also includes a parser that determines a context, selects a domain agent based on that context, and transforms the user's words into a command formulated in a grammar used by the selected agent.
- Finally, the system includes an agent architecture that couples the services of an agent manager, system agent, and domain agents to create and format a response.
U.S. Patent No. 8,015,006 - “Systems And Methods For Processing Natural Language Speech Utterances With Context-Specific Domain Agents”
- Issued: September 6, 2011 (Compl. ¶49).
The Invention Explained
- Problem Addressed: The patent identifies the difficulty of machine communication with humans, noting that machine-based queries are often "highly structured and are not inherently natural to the human user" (’006 Patent, col. 1:38-41; Compl. ¶51). It specifically highlights that most natural language queries are incomplete, ambiguous, or subjective, which poses a "significant barrier" to human-machine interaction (Compl. ¶51).
- The Patented Solution: The invention proposes a method that "makes significant use of context, prior information, domain knowledge, and user specific profile data" to enable a more natural conversational environment (’006 Patent, abstract; Compl. ¶53). The system uses "domain agents" to manage domain-specific information and behavior (’006 Patent, col. 2:53-3:7; Compl. ¶53). A key part of the process involves receiving a speech utterance, recognizing the words, parsing the information to determine meaning and context, and then using a domain agent-specific grammar to formulate, process, and respond to the user's request (’006 Patent, fig. 1; Compl. ¶54).
- Technical Importance: This technology provided a framework for processing imprecise human speech by using context and domain-specific modules (agents) to infer user intent and formulate a structured, machine-executable request (Compl. ¶52; Compl. ¶53).
Key Claims at a Glance
- The complaint asserts independent Claim 5 (Compl. ¶137).
- The essential elements of Claim 5, a method, include:
- Receiving a natural language speech utterance containing a request.
- Recognizing words or phrases from the utterance using dictionary and phrase tables.
- Parsing the utterance to determine a meaning and a context.
- Formulating the request in accordance with a grammar used by a domain agent associated with the determined context.
- This "formulating" step includes sub-steps of determining required/optional values, extracting criteria/parameters, inferring further criteria/parameters using a dynamic set of prior probabilities or fuzzy possibilities, and transforming these into tokens compatible with the agent's grammar.
- Processing the formulated request with the domain agent to generate a response.
- Presenting the response via the speech unit.
U.S. Patent No. 8,140,327 - “System And Method For Filtering And Eliminating Noise From Natural Language Utterances To Improve Speech Recognition And Parsing”
- Patent Identification: U.S. Patent No. 8,140,327, issued March 20, 2012 (Compl. ¶58).
- Technology Synopsis: The patent addresses the problem of background noise in speech recognition systems (Compl. ¶60). The solution combines a microphone array that can create "nulls" to notch out noise sources, an adaptive filter that uses various techniques including echo cancellation, a speech coder that uses adaptive lossy compression, and a transceiver that communicates the digitized signal at a rate dependent on available bandwidth (Compl. ¶61).
- Asserted Claims: The complaint asserts independent Claim 14 (Compl. ¶178).
- Accused Features: The complaint alleges that Alexa devices, particularly technology found in the "Amazon Alexa Premium Voice Far-Field Development Kit," infringe by using microphone arrays, beamforming to notch out noise, and adaptive filters for echo cancellation (Compl. ¶¶180-182; Compl. ¶184).
U.S. Patent No. 8,195,468 - “Mobile Systems And Methods Of Supporting Natural Language Human-Machine Interactions”
- Patent Identification: U.S. Patent No. 8,195,468, issued June 5, 2012 (Compl. ¶65).
- Technology Synopsis: The patent addresses the problem that verbal communications are "fundamentally incompatible" with machine processing because user speech is natural while machine requests must be highly structured (Compl. ¶68). The invention describes a method for processing multi-modal (speech and non-speech) inputs by creating and merging transcriptions of both, using a semantic knowledge-based model (including personalized, general, and environmental models) to determine the most likely context from a "context stack," and then identifying a domain agent to process the request (Compl. ¶69).
- Asserted Claims: The complaint asserts independent Claim 19 (Compl. ¶213).
- Accused Features: The complaint alleges that multi-modal Alexa products like the Echo Show infringe by receiving and processing both speech and non-speech (e.g., screen-based) inputs, using user identity and prior interactions to personalize responses (Compl. ¶¶215-216; Compl. ¶221).
U.S. Patent No. 9,031,845 - “Mobile Systems And Methods For Responding To Natural Language Speech Utterance”
- Patent Identification: U.S. Patent No. 9,031,845, issued May 12, 2015 (Compl. ¶73).
- Technology Synopsis: The patent addresses the challenges of creating a speech interface for a vehicular environment, including the difficulty of processing commands that may need to be executed either locally on the vehicle or remotely via a wireless network (Compl. ¶¶76-77). The solution is a mobile system that receives a natural language utterance, determines its domain and context, formulates a command, and then determines whether to execute that command on-board the vehicle or off-board via a wireless device (Compl. ¶79).
- Asserted Claims: The complaint asserts independent Claim 1 (Compl. ¶251).
- Accused Features: The complaint alleges that Amazon’s automotive products, such as the Alexa Auto SDK, infringe by providing for both on-board command execution (via a "Local Voice Control" extension) and off-board execution via the Alexa cloud (Compl. ¶¶264-265; Compl. ¶269).
U.S. Patent No. 9,263,039 - “Systems And Methods For Responding To Natural Language Speech Utterance”
- Patent Identification: U.S. Patent No. 9,263,039, issued February 16, 2016 (Compl. ¶83).
- Technology Synopsis: The patent addresses the problem of processing imperfect user communications, such as incomplete phrases or slang (Compl. ¶86). The invention describes a method for processing both speech and non-speech communications by transcribing and merging them, searching the merged query for text combinations, comparing those combinations to a "context description grammar," generating a relevance score, selecting one or more domain agents based on that score, and generating a response (Compl. ¶87).
- Asserted Claims: The complaint asserts independent Claim 13 (Compl. ¶290).
- Accused Features: The complaint alleges that multi-modal Alexa products like the Echo Show infringe by processing both speech and non-speech inputs, comparing the inputs against grammars associated with skills, and using a relevance scoring or ranking system (e.g., "Shortlister and HypRank") to select the most relevant skill to invoke (Compl. ¶¶292; Compl. ¶302; Compl. ¶¶305-308).
U.S. Patent No. 9,495,957 - “Mobile Systems And Methods Of Supporting Natural Language Human-Machine Interactions”
- Patent Identification: U.S. Patent No. 9,495,957, issued November 15, 2016 (Compl. ¶91).
- Technology Synopsis: The patent addresses the difficulty of interpreting natural language utterances, which are often incomplete or ambiguous (Compl. ¶93). The solution is a system that processes a natural language utterance by generating a "context stack" from prior utterances, performing speech recognition on the new utterance, comparing the recognized words to entries in the context stack, generating rank scores for the context entries, and using the highest-ranked entries to determine the user's command or request (Compl. ¶95).
- Asserted Claims: The complaint asserts independent Claim 1 (Compl. ¶331).
- Accused Features: The complaint alleges that Alexa devices infringe by using a "shortlisting-reranking approach" that uses a record of past interactions (a "context stack") to interpret a new natural language request and identify the most relevant skills based on ranked scores (Compl. ¶¶334-335; Compl. ¶¶342-343).
III. The Accused Instrumentality
Product Identification
- The complaint collectively refers to the accused instrumentalities as the “Alexa Products” (Compl. ¶2). This includes Amazon’s Alexa virtual assistant, the Echo line of smart speakers and displays (e.g., Echo, Echo Dot, Echo Show), Amazon’s mobile applications (e.g., Alexa app, Music app), the Alexa cloud infrastructure, and Alexa Voice Services (AVS) for integration into third-party devices (Compl. ¶98). The complaint also specifically identifies automotive implementations such as Echo Auto and the Alexa Auto SDK (Compl. ¶103).
Functionality and Market Context
- The accused products provide a voice-controlled user interface for a wide array of consumer electronics and services (Compl. ¶1). A user speaks a "wake word" followed by a natural language command or question (Compl. ¶110). This utterance is processed by Amazon’s system, which includes Automatic Speech Recognition (ASR) to convert speech to text and Natural Language Understanding (NLU) to determine the user’s intent (Compl. ¶106). The system then identifies and invokes an appropriate application, which the complaint refers to as a "skill" or "domain agent," to fulfill the request (Compl. ¶110; Compl. ¶114). A diagram in the complaint illustrates this workflow, showing a user request being processed by the Alexa service (ASR, NLU) which then sends a structured request to the appropriate "skill logic" for execution (Compl. p. 65). The system is alleged to use conversational history to inform its interpretation of new utterances (Compl. ¶112).
IV. Analysis of Infringement Allegations
U.S. Patent No. 7,693,720 Infringement Allegations
| Claim Element (from Independent Claim 1) | Alleged Infringing Functionality - | Complaint Citation | Patent Citation |
| a speech unit connected to a computer device on a vehicle, wherein the speech unit receives a natural language speech utterance from a user and converts the received natural language speech utterance into an electronic signal | The Accused Automotive Products and Services include in-vehicle hardware, such as an in-cabin microphone and speakers connected to an automotive head unit, which receives a user's speech and converts it into an electronic signal. - | ¶107; ¶108 | col. 3:1-13 |
| a speech recognition engine that recognizes at least one of words or phrases...wherein the data used by the speech recognition engine includes a plurality of dictionary and phrase entries that are dynamically updated based on at least a history of a current dialog and one or more prior dialogs associated with the user | The Alexa system is alleged to recognize words and phrases using information from different intents and to track "previously provided information," maintaining a history of interactions "to learn more about you as they listen" and help build new Alexa experiences. | ¶111; ¶112; Compl. p. 68 | col. 27:1-24 |
| a parser that interprets the recognized words or phrases by...determining a context for the natural language speech utterance; selecting at least one of the plurality of domain agents based on the determined context; and transforming the recognized words or phrases into at least one of a question or a command...formulated in a grammar that the selected domain agent uses | The Alexa system allegedly interprets recognized words to determine user intent (context) and selects the most relevant "skill" (domain agent) to handle the request. The system formulates a question or command based on the grammar of different Alexa intents. | ¶113; ¶114 | col. 28:5-41 |
| an agent architecture that communicatively couples services of...an agent manager, a system agent, the plurality of domain agents, and an agent library...wherein the selected domain agent uses the communicatively coupled services to create a response...and format the response | The Alexa platform is described as providing the infrastructure to support a variety of "skills" (domain agents), which are analogous to apps. The selected skill interacts with the user and the broader Alexa infrastructure to generate and present a response. | ¶115; ¶116 | col. 21:61-22:57 |
U.S. Patent No. 8,015,006 Infringement Allegations
| Claim Element (from Independent Claim 5) | Alleged Infringing Functionality - | Complaint Citation | Patent Citation |
| receiving, at a speech unit coupled to a processing device, a natural language speech utterance that contains a request | The accused Alexa devices, such as an Echo Dot or Echo Show, receive a user request in the form of a natural language utterance. - | ¶140; ¶141 | col. 23:25-34 |
| recognizing, at a speech recognition engine..., one or more words or phrases contained in the utterance using information in one or more dictionary and phrase tables | Alexa processes the user request with speech recognition and natural language understanding, using a wide range of sentences, phrases, and words. - | ¶142; ¶143 | col. 16:1-14 |
| parsing, at a parser..., information relating to the utterance to determine a meaning associated with the utterance and a context associated with the request | Alexa allegedly uses "intents" to represent the action or meaning of a request, and "slots" to capture variable information that provides context for the intent. A diagram in the complaint shows how information such as "San Francisco" can be assigned to different slots (e.g., WeatherLocation, City, Town) to maintain context across a conversation. | ¶144; ¶145; Compl. p. 81 | col. 17:13-40 |
| inferring one or more further criteria and one or more further parameters associated with the request using a dynamic set of prior probabilities or fuzzy possibilities | Alexa is alleged to infer parameters for a request by using the history of the current interaction. The complaint alleges Amazon's approach makes decisions "about slot values mentioned in context" and "the probability that any given carryover decision is the correct one." | ¶152; ¶153; Compl. p. 86 | col. 3:19-35 |
| processing the formulated request with the domain agent associated with the determined context to generate a response to the utterance | Alexa is alleged to process the request using a selected "skill" (domain agent), which includes a "skill interaction model" and "skill application logic" to produce a response. - | ¶156; ¶157 | col. 18:50-19:12 |
- Identified Points of Contention:
- Scope Questions: The ’720 Patent claims are directed to a "mobile system... on a vehicle" (Compl. ¶45). A central question may be whether Amazon's general Alexa platform, which is also used in non-mobile, home environments, can be said to practice these claims, or if the infringement analysis is limited only to the specific automotive products like Echo Auto.
- Technical Questions: A key technical question will likely concern the nature of the "domain agents" and the "dynamic set of prior probabilities." The patents were filed in the early 2000s and describe a specific agent-based architecture (’006 Patent, fig. 6). The infringement analysis may turn on whether Amazon's modern "skills" ecosystem and its use of neural networks and machine learning models function in a way that is equivalent to the claimed "domain agents" and the "dynamic set of prior probabilities or fuzzy possibilities" as understood by a person of ordinary skill in the art at the time of the invention.
V. Key Claim Terms for Construction
The Term: "domain agent" (’720 Patent, Claim 1; ’006 Patent, Claim 5)
- Context and Importance: This term is foundational to the architecture of the asserted inventions. The complaint consistently equates Amazon's "skills" with the claimed "domain agents" (Compl. ¶110; Compl. ¶114; Compl. ¶147). The viability of the infringement case may depend heavily on whether the court construes "domain agent" broadly enough to read on the Alexa skills architecture.
- Intrinsic Evidence for Interpretation:
- Evidence for a Broader Interpretation: The specification describes agents as "executables that receive, process and respond to user questions, queries and commands" and as "re-distributable packages or modules of functionality, typically for a specific domain" (’720 Patent, col. 4:5-11). This functional description could support a broad reading that encompasses Amazon's skills.
- Evidence for a Narrower Interpretation: The specification also provides a specific diagram of an "Agent Architecture" that includes an "agent manager," a "system agent," and an "agent library" with "criteria handlers" (’720 Patent, fig. 6; ’720 Patent, col. 22:1-12). This could support a narrower construction limited to systems that share this specific architectural structure, which Amazon may argue its Alexa Skills Kit does not.
The Term: "dynamically updated" (referring to "dictionary and phrase entries") (’720 Patent, Claim 1)
- Context and Importance: This limitation distinguishes the invention from static, predefined command systems. The infringement allegation hinges on whether Alexa's method of learning from user interactions and conversational history meets this definition (Compl. ¶112). Practitioners may focus on whether "dynamically updating" a "dictionary" requires a specific technical mechanism (e.g., adding a string to a list) or can be read more broadly to cover the evolving state of a machine learning model.
- Intrinsic Evidence for Interpretation:
- Evidence for a Broader Interpretation: The patent describes the goal as interpreting questions "in the context of previous questions, knowledge of the domain, or the user's history of interests and preferences" (’720 Patent, col. 2:5-8). This purpose-driven language could support interpreting "dynamically updated" to mean any method of adapting to new information from the user's history.
- Evidence for a Narrower Interpretation: The specification describes a process where a user can spell an unrecognized word, and its "pronunciation" is then "added to either the dictionary, the agent 106, or the user's profile 110" (’006 Patent, col. 27:45-50, a related patent). This suggests a more literal mechanism of adding discrete entries to a stored list, which could support a narrower construction.
VI. Other Allegations
- Indirect Infringement: The complaint alleges inducement of infringement, stating that Amazon knowingly encourages and instructs businesses and consumers to use the Accused Products in their ordinary and intended infringing manner (Compl. ¶122; Compl. ¶165). It also alleges contributory infringement, stating Amazon provides components that are a material part of the invention and are not staple articles of commerce suitable for substantial non-infringing use (Compl. ¶124; Compl. ¶167).
- Willful Infringement: The complaint alleges willful infringement based on Amazon’s alleged pre-suit knowledge of the patents (Compl. ¶¶130; Compl. ¶171). The basis for this knowledge is a series of communications, presentations, and licensing negotiations between the original patent owner, VoiceBox, and Amazon, which allegedly began in 2011 and included presentations detailing VoiceBox's "patented technology" (Compl. ¶¶28-31; Compl. ¶118). The complaint also points to Amazon's subsequent hiring of key VoiceBox technical personnel as further evidence of knowledge (Compl. ¶35; Compl. ¶118).
VII. Analyst’s Conclusion: Key Questions for the Case
- A core issue will be one of definitional scope: can the term "domain agent," described in the context of a specific agent-based architecture from the early 2000s, be construed to cover applications in Amazon’s modern, cloud-based "skills" ecosystem? This will likely involve a detailed comparison of the patent's described architecture with the functionality of the Alexa Skills Kit.
- A second central question will be one of technical mechanism: does the process by which Amazon's machine learning models adapt and learn from conversational history constitute the claimed method of "dynamically updat[ing] a plurality of dictionary and phrase entries," or is there a fundamental operational difference between updating a discrete list and refining a statistical model?
- A key evidentiary question will concern pre-suit knowledge: what specific technical details regarding the patented inventions were disclosed to Amazon during the 2011 meetings with VoiceBox, and can the plaintiff demonstrate that this disclosure was sufficient to establish that Amazon knew or should have known that its subsequent development of the Alexa platform would infringe the specific claims now at issue?