

Multimodal embedding model that creates holistic video/audio/text/image embeddings for semantic search and video understanding.

Multimodal embedding model that creates holistic video/audio/text/image embeddings for semantic search and video understanding.
Marengo 3.0 is TwelveLabs' multimodal embedding model designed to analyze and encode visuals, audio, and text from videos into dense embeddings. It enables semantic video search, content indexing, and downstream tasks by producing high-quality, multimodal representations that capture scene, audio, and textual context. The model is accessible via TwelveLabs SDKs (Python and JavaScript) and is used in embedding pipelines (including asynchronous workflows on platforms like AWS Bedrock) to create searchable indexes, frame- and shot-level analyses, and transcript-linked embeddings. Its value lies in providing holistic, human-like video understanding that supports image- and natural-language queries for efficient retrieval and analysis.


TwelveLabs Marengo 3.0 features a flexible pricing structure that includes a Free plan, an Enterprise/API access option with custom pricing, a monthly service plan, and usage-based pricing available through Amazon Bedrock. For specific rates and tailored solutions, it's best to contact TwelveLabs directly.
TwelveLabs Marengo 3.0 offers multiple pricing tiers to accommodate various user needs:
Free Plan: Ideal for individuals or small projects, this plan provides essential functionalities without any cost. It's an excellent way to explore the platform's capabilities before committing financially.
Enterprise/API Access: This option is tailored for organizations requiring extensive features or integration capabilities. Pricing is customized based on specific needs, such as the volume of data processed or the complexity of API calls. Businesses should reach out to TwelveLabs for a personalized quote.
Monthly Service Plan: This plan offers users a straightforward monthly payment structure, making it easier for businesses to budget their expenses. It’s suitable for medium-sized organizations that benefit from consistent access to the platform's features.
Usage-Based Pricing: Available through Amazon Bedrock, this pricing model allows users to pay only for what they use. This is beneficial for those who may not require constant access but need flexibility during peak times. Users can expect variable costs depending on their consumption patterns.
TwelveLabs Marengo 3.0 provides advanced video understanding features through multimodal embedding that integrates video, audio, text, and images. Key capabilities include frame-level analysis, customizable embedding parameters, and software development kits (SDKs) for seamless integration, enabling enhanced semantic search and improved content discovery.
TwelveLabs Marengo 3.0 stands out in the field of video understanding by utilizing multimodal embedding. This technology enables the system to analyze and correlate various types of media, such as video clips, audio tracks, textual information, and images, thereby creating a richer understanding of the content.
Multimodal Embedding: By combining multiple data types, users can perform semantic searches that yield more relevant results. For instance, a video search may return results that include related audio commentary or contextual images.
Frame-Level Analysis: This feature allows for the extraction of insights from individual frames within videos. For example, a security application could leverage frame-level analysis to detect specific actions or events, enhancing real-time monitoring capabilities.
Integration SDKs: The Marengo 3.0 SDKs enable developers to incorporate these advanced features into their applications. This means businesses can tailor the functionality to meet their specific needs, whether they are in media production, security, or marketing.
Configurable Embedding Parameters: Users can adjust embedding settings to optimize performance for specific use cases, such as enhancing accuracy in video search results or reducing processing time for real-time applications.
To start using TwelveLabs Marengo 3.0, sign up for the Free plan on the TwelveLabs platform. Then, explore the SDKs for Python and JavaScript to create indexes and perform embedding tasks efficiently. This allows you to harness the power of AI in your applications seamlessly.
Getting started with TwelveLabs Marengo 3.0 is straightforward. First, visit the TwelveLabs website and create an account by selecting the Free plan. This plan provides access to essential features for users to explore the platform's capabilities without any financial commitment.
Once your account is set up, download the SDKs for either Python or JavaScript. These SDKs facilitate the integration of TwelveLabs features into your applications. For instance, in Python, you can easily create an index using just a few lines of code. Here's a simple example:
from twelve_labs import TwelveLabs
# Initialize the TwelveLabs client
client = TwelveLabs(api_key='YOUR_API_KEY')
# Create an index
index = client.create_index(name='my_index')
After creating your index, you can start performing embedding tasks. This process involves converting text or other data types into a numerical format that machine learning algorithms can understand. The TwelveLabs documentation provides comprehensive guides and code snippets to assist you along the way.
Yes, TwelveLabs Marengo 3.0 seamlessly integrates with AWS services, particularly through Amazon Bedrock. This integration enables efficient cloud-based embedding generation and storage solutions, such as Amazon S3, enhancing data management and deployment capabilities for developers and businesses.
TwelveLabs Marengo 3.0 provides robust integration with AWS, making it a powerful tool for developers looking to enhance their AI and machine learning capabilities. By connecting with Amazon Bedrock, Marengo 3.0 allows users to deploy AI models with ease, utilizing pre-trained models and the ability to customize them for specific applications.
TwelveLabs Marengo 3.0 excels in comparison to alternative AI models for video analysis due to its advanced multimodal embedding capabilities. This technology enables comprehensive analysis across diverse media formats, significantly improving semantic search, video comprehension, and contextual understanding.
TwelveLabs Marengo 3.0 utilizes cutting-edge multimodal embedding technology, which allows it to analyze and integrate information from different types of media—such as audio, text, and video. This capability enables the model to create a unified representation of content, leading to superior understanding and interpretation.
For instance, while traditional models may focus solely on video frames, Marengo 3.0 also considers accompanying audio transcripts and metadata. This holistic approach allows it to recognize context more effectively, improving the accuracy of video searches. In practical applications, this means that users can find specific scenes or moments in videos based on nuanced queries, such as searching for "scenes with emotional dialogues."
Moreover, Marengo 3.0's performance is backed by advanced machine learning algorithms that continuously refine its capabilities through user interactions and feedback. This leads to a more adaptive system that evolves as it processes more data.
Browse by use case: Video Generation · Voice & Audio
Compare TwelveLabs Marengo 3.0: vs Laguna by Poolside · vs Arena AI: The Official AI Ranking & LLM Leaderboard · vs PromptLayer · vs PHBench