DCT

1:26-cv-07481

BTF IP Holdings LLC v. Eleven Labs Inc

Key Events
Complaint
complaint Intelligence

I. Executive Summary and Procedural Information

  • Parties & Counsel:
  • Case Identification: 1:26-cv-07481, S.D.N.Y., 09/01/2026
  • Venue Allegations: Venue is alleged to be proper in the Southern District of New York because Defendant maintains its headquarters and a regular and established place of business in the district, and has placed its services into the stream of commerce there.
  • Core Dispute: Plaintiff alleges that Defendant's AI-powered voice generation platform infringes a patent related to methods for embodying an online media service with a multiple voice system.
  • Technical Context: The technology at issue involves AI-driven text-to-speech, voice cloning, and audio content generation, a rapidly growing field central to digital media, content creation, and accessibility services.
  • Key Procedural History: The complaint alleges that Plaintiff notified Defendant of its infringing activities at least as early as July 2026. No other procedural events, such as prior litigation or administrative proceedings involving the patent, are mentioned.

Case Timeline

Date Event
2019-09-18 '593 Patent - Earliest Priority Date
2022-01-01 Defendant Eleven Labs Inc. Founded
2022-12-06 '593 Patent - Issue Date
2026-07-01 Alleged Notification of Infringement to Defendant
2026-09-01 Complaint Filing Date

II. Technology and Patent(s)-in-Suit Analysis

  • Patent Identification: U.S. Patent No. 11,521,593 ("the '593 Patent"), "METHOD OF EMBODYING ONLINE MEDIA SERVICE HAVING MULTIPLE VOICE SYSTEMS," issued December 6, 2022.
  • The Invention Explained:
    • Problem Addressed: The patent's background section identifies the limitations of consuming text-based online content, particularly on mobile devices or while engaged in other activities like driving ʼ593 Patent, col. 1:46-56 It describes conventional text-to-speech technology as "boring" and lacking a "system for participation of readers" ʼ593 Patent, col. 2:1-6
    • The Patented Solution: The invention describes a method to make audio consumption of online articles more dynamic and personalized. The system collects content from a media site, allows a subscriber to either input their own voice or select a pre-stored voice (which can be bought and sold in an "online store"), classifies the content, converts it to speech using the selected voice, and outputs the result ʼ593 Patent, abstract The patent also discloses adding background sounds or specific intonations to the audio output ʼ593 Patent, col. 6:25-39 ʼ593 Patent, Fig. 7
    • Technical Importance: The described method aims to transform passive, machine-driven text-to-speech into an interactive and customizable experience, thereby increasing user engagement with online content ʼ593 Patent, col. 2:7-19
  • Key Claims at a Glance:
    • The complaint asserts infringement of at least Claim 1 of the '593 Patent Compl. ¶14
    • Independent Claim 1 requires a method comprising:
      • A first operation of collecting and displaying online content from a media site.
      • A second operation of inputting a subscriber's voice or selecting a pre-stored voice, where the voice can be converted into a selected language.
      • A third operation of recognizing and classifying the online content.
      • A fourth operation of converting the classified content into speech.
      • A fifth operation of outputting the speech using the selected voice.
      • The claim further requires that the second operation includes pre-storing, purchasing, and selling voices in an online store.
      • The claim also requires that the fifth operation includes selecting online content, selecting different background sounds for different sections of the content, and outputting the content with the selected voice and background sounds.

III. The Accused Instrumentality

  • Product Identification: The accused instrumentality is Defendant's "voice generation platform" ("VGP"), also referred to as the "Accused Service" Compl. ¶2
  • Functionality and Market Context: The Accused Service is described as a platform that "transforms text into lifelike speech" and allows users to generate voices for videos, podcasts, and other media Compl. ¶2 It is accessible via a web application, a REST API, and SDKs Compl. ¶15 Key alleged functionalities include text-to-speech conversion, voice cloning from audio samples, a "Voice Library" of pre-stored voices, and the ability to dub content into different languages Compl. ¶¶15, 19, 22 The complaint presents a screenshot from Defendant's website illustrating the platform's capabilities, including text-to-speech, voice cloning, and conversational agents Compl. ¶15, Ex. B

IV. Analysis of Infringement Allegations

'593 Patent Infringement Allegations

Claim Element (from Independent Claim 1) Alleged Infringing Functionality Complaint Citation Patent Citation
a first operation of [1.1.1] collecting preset online articles and content from a specific media site and [1.1.2] displaying the online articles and content on a screen of a personal terminal The VGP collects content from a specific media site by allowing a user to import a URL, which is then displayed on the user's terminal. A screenshot shows the 'Import URL' dialog box Compl. ¶17, Ex. D ¶17; ¶18 col. 4:12-17
a second operation of inputting a voice of a subscriber or setting a voice of a specific person among voices that are pre-stored in a database The VGP allows a subscriber to input their voice via "voice cloning options" or, alternatively, to select a pre-stored voice from a database. A screenshot displays options for "Instant Voice Clone" and "Professional Voice Clone" Compl. ¶19, Ex. F ¶19; ¶20 col. 4:17-20
wherein the second operation further comprises: [1.2.3] pre-storing the voice of the specific person in an online store for each field; [1.2.4] purchasing the voice...; and [1.2.5] directly registering and selling the voice of the subscriber in the online store The complaint alleges that selecting a paid subscription plan constitutes "purchasing" the option to use a cloned or specific voice. It further alleges voices are "pre-stored" and "sold" via the "Voice Library," which it equates to an "online store." A screenshot of subscription tiers is provided as evidence Compl. ¶31, Ex. S ¶31; ¶32 col. 4:39-45
a third operation of recognizing and classifying the online articles and content The VGP allegedly recognizes and classifies online articles by distinguishing them from other content like advertisements and sidebars when importing from a URL. ¶25 col. 4:20-21
a fourth operation of converting the classified online articles and content into speech The VGP converts the imported and classified text content into audio speech, for instance, by clicking a "Generate" button. ¶26; ¶27 col. 4:21-23
a fifth operation of outputting the speech obtained in the fourth operation using the voice of the subscriber or the specific person The VGP outputs the generated speech as an audio file using the voice previously cloned by the subscriber or selected from the library. ¶28; ¶29; ¶30 col. 4:23-26
wherein the fifth operation further comprises: [1.5.2] selecting different types of background sounds according to different sections of the online articles and content The complaint alleges the VGP allows for the selection of different background sounds, which are added to a project timeline. A screenshot shows a "Sound Effects" menu with various audio options available to be added to the project Compl. ¶34, Ex. V ¶34 col. 6:25-39
  • Identified Points of Contention:
    • Scope Questions: A potential point of contention is whether the accused user-driven action of importing a URL Compl. ¶17 meets the claim limitation of "collecting preset online articles." The term "preset" may suggest that the service curates or pre-selects the content, which could create a scope mismatch with the accused functionality.
    • Technical Questions: The complaint alleges that the VGP's ability to add "Sound Effects" to a timeline Compl. ¶34 satisfies the limitation of "selecting different types of background sounds according to different sections of the online articles and content." A question for the court may be whether the general addition of sound effects to a project timeline is technically equivalent to selecting sounds based on "different sections" of the content, as the complaint does not provide evidence of this specific link.

V. Key Claim Terms for Construction

  • The Term: "online store"

  • Context and Importance: Claim 1 requires several actions related to an "online store," including pre-storing, purchasing, and selling voices. The complaint maps this term to the Defendant's "Voice Library" and its subscription-based business model Compl. ¶¶31-32 The viability of the infringement allegation hinges on whether this model, where users pay a subscription fee for access and features, constitutes "purchasing" and "selling" voices in an "online store" as contemplated by the patent.

  • Intrinsic Evidence for Interpretation:

    • Evidence for a Broader Interpretation: The specification describes a system for making voices available for a fee, stating a subscriber can "purchase" a voice at an "online store" ʼ593 Patent, col. 6:10-14 This could be argued to broadly cover any commercial mechanism for obtaining access to voices, including a subscription.
    • Evidence for a Narrower Interpretation: The patent discusses "registering and selling the voice of the subscriber in the online store" and setting a "desired selling price" ʼ593 Patent, col. 6:15-23 ʼ593 Patent, Fig. 6 This language, along with Figure 6 which depicts a transactional flow for registering and selling a specific voice, may support a narrower interpretation of a direct, per-voice marketplace rather than a platform-wide subscription service.
  • The Term: "recognizing and classifying the online articles and content"

  • Context and Importance: This term is central to the data processing aspect of the invention. The complaint alleges that the VGP's ability to ignore advertisements and sidebars when importing a URL meets this limitation Compl. ¶25 Practitioners may focus on this term because its scope will determine whether simple content extraction is sufficient to infringe, or if a more sophisticated analysis (e.g., by topic, sentiment, or structure) is required.

  • Intrinsic Evidence for Interpretation:

    • Evidence for a Broader Interpretation: The patent states that "an image or text region may be extracted" and that content can be classified "for each title or content" ('593 Patent, col. 4:47-54). This could be interpreted broadly to cover any act of identifying and isolating the main body of an article.
    • Evidence for a Narrower Interpretation: The specification also describes classifying content based on "section, keyword, article, news agency, latest news article, date, view count, degree of association, or headline" ('593 Patent, col. 4:56-61). This detailed list suggests a more granular classification may be required than simply separating an article from advertisements.

VI. Other Allegations

  • Indirect Infringement: The complaint alleges inducement by asserting that Defendant encourages and facilitates infringement by marketing and providing the Accused Service with instructions on how to use its features Compl. ¶13 It also alleges contributory infringement, stating the Accused Service is not a staple article of commerce and is especially adapted for infringing the '593 Patent Compl. ¶13
  • Willful Infringement: The claim for willful infringement is based on the allegation that Defendant had knowledge of the '593 Patent and its infringement since at least July 2026, following direct notification from the Plaintiff Compl. ¶36

VII. Analyst's Conclusion: Key Questions for the Case

  • A core issue will be one of definitional scope: does the term "collecting preset online articles," which may imply a curated or automated feed, read on the accused functionality where a user manually provides a URL or text for conversion?
  • A second key issue will turn on claim construction: can the term "online store," which the patent illustrates with direct transactional language of buying and selling individual voices, be construed to cover Defendant's subscription-based model that provides access to a library of voices and voice cloning features?
  • A central evidentiary question will be one of functional equivalency: does the Accused Service's feature for adding generic "Sound Effects" to a project timeline perform the specific function required by Claim 1 of "selecting different types of background sounds according to different sections of the online articles and content," or does a technical mismatch exist in their operation?