Topic

Bright Data

All digests tagged Bright Data

Physical AI's Next Bottleneck Is Finding the Right Video — Rafael Levi, Bright Data thumbnail

· 16:32

Physical AI's Next Bottleneck Is Finding the Right Video — Rafael Levi, Bright Data

The primary bottleneck in advanced AI development is no longer the model architecture, but the availability and quality of training data. While Large Language Models (LLMs) have access to trillions of words, robotics training is severely limited, relying on small, often biased datasets (e.g., only about a million robot videos). The speaker, Rafael Levi of Bright Data, argues that the solution lies in leveraging the massive, real-world video data available on the web. Bright Data's approach, 'search first, collect second,' indexes billions of videos by specific actions, allowing users to retrieve precisely trimmed, ready-to-use video snippets via an API, drastically reducing data noise and collection waste compared to traditional methods.

Key takeaways

  1. The Data Bottleneck in Robotics 4:00

    For robotics, training data is scarce compared to text (trillions of words) or images (billions of labeled images), with current datasets limited to about a million videos of robots doing actions. (2:40)

  2. Bias in Staged Data 3:15

    Paying people to record staged actions (e.g., opening a door) results in biased data because human behavior is different when uninstructed versus when performing a task for recording. (3:15)

  3. The Web as a Solution 5:14

    The web holds billions of hours of real-world video showing natural physics, gravity, and cause-and-effect interactions, making it a massive training source for world models. (5:14)

  4. Bright Data's Indexing Approach 8:33

    Instead of downloading massive amounts of video and discarding the majority (e.g., NVIDIA discarding 96% of downloaded video), Bright Data indexes videos by specific actions, allowing users to search for precise clips (e.g., 'washing dishes,' 'folding a t-shirt'). (8:33)

Watch on YouTube Full article

The Rise of CaaS: Context-as-a-Service for Agentic AI — Omer Primor, Bright Data thumbnail

· 22:20

The Rise of CaaS: Context-as-a-Service for Agentic AI — Omer Primor, Bright Data

The video analyzes the shift from viewing web data as a simple source of information to treating it as dynamic 'context' for agentic AI. The speaker argues that Context-as-a-Service (CaaS) vendors are emerging to provide structured knowledge graphs, acting as vertical search engines. Critically, he emphasizes that at scale, the cost killer is not initial volume but the *frequency* of repeated queries. For persistent knowledge work, owning and building a custom data pipeline—even if time-consuming—can eventually become more cost-effective than continually renting context from third-party vendors.

Key takeaways

  1. Context Decay: Data is never a snapshot 0:02

    Web data decays quickly (e.g., social content < 1 day; news/finance ~30 days). Therefore, extracting context must be treated as an ongoing process, not a one-time effort [2:43].

  2. The Rise of CaaS for Agents 0:06

    AI agents require structured knowledge beyond what general search provides. CaaS vendors address this by developing and indexing specialized knowledge graphs (vertical search) across multiple data sources, enabling deep reasoning [6:32].

  3. Frequency is the Cost Killer at Scale 0:12

    When performing repeated due diligence or market research, every query costs money, even if nothing has changed. This recurring cost (frequency) eventually surpasses the initial setup cost of building an owned pipeline [12:32].

  4. The Tipping Point for Ownership 0:15

    There is a tipping point where the cumulative cost of repeated context queries makes it economically viable to build and own the data retrieval pipeline in-house, potentially bypassing middleman costs [15:22].

Watch on YouTube Full article