
Blog written by Abhijit Mugali, Product Manager for Test Data Manager
Who will capture the true benefits of Generative AI (GenAI) first?
Organizations are racing to innovate–driving faster software development, better end-user experiences, differentiated products, enhanced automation, and more personalized customer interactions. To achieve these goals, many are turning to Generative AI as a key enabler of transformation and competitive advantage.
As organizations increasingly adopt GenAI models, the need for efficient and scalable testing practices becomes more critical than ever. Tuning and testing these models requires vast amounts of diverse, high-quality test data to ensure accuracy, reliability, and performance.
With Broadcom’s DAVE (Database Virtualization Engine), organizations gain access to the high-quality and high-volume test data they need to effectively tune and test GenAI models. DAVE simplifies the traditionally complex, resource-intensive, and costly task of provisioning test data, making it faster and more efficient for teams to support GenAI development and testing.
With the data available on demand, the designing, tuning, and testing cycles become faster and more efficient than ever before. Are you looking to improve the overall quality of AI-driven applications? If so, DAVE can help.
Getting to Know DAVE
DAVE (Database Virtualization Engine) is a part of Broadcom’s overall data management platform. Database virtualization is a technology that abstracts the underlying complexity of databases and allows multiple, isolated, virtual databases to be created from a single physical database. This capability enables organizations to rapidly provision environments without the need for dedicated physical resources or dealing with intricate schema dependencies.
In modern testing environments, especially for complex AI models, Broadcom’s data platform along with DAVE allows testers to quickly access secure, compliant, realistic, production-like data without the need to replicate entire physical datasets. This capability streamlines the test data engineering process: improving lead time to change, accelerating testing efforts, and improving overall quality. This is particularly valuable for embedding agile development and DevOps practices in AI-driven projects in order to improve time to market and business outcomes.
Modernizing GenAI Models with DAVE
Testing GenAI models requires a myriad of datasets. Datasets are required to validate AI models and are needed to accommodate varied scenarios and ensure accurate predictions are being made. One of the challenges is generating enough variations of the data to properly train and evaluate these models. DAVE makes it easy for teams to create and provision multiple virtual copies of their datasets. This simplifies dataset versioning, managing variations, data creation, and environment provisioning. It also helps reduce overall storage and environment costs. As a result, GenAI models can be tested with a wide range of inputs, leading to better results and outcomes, without increasing environment complexity or costs.
What are the key ways DAVE helps GenAI Initiatives:
-
Quickly Provisioning Diverse Datasets GenAI models thrive on large volumes of data. Traditional methods of data provisioning can be time-consuming, especially when data relationships are complex or schema definitions are unclear. With DAVE, teams can quickly generate multiple versions of the same data or apply variations to build robust data sets. Having data on demand with different inputs and data combinations, allows teams to rapidly develop, augment, tune, and test how the model performs, speeding up the development cycle.
-
Simulating Real-World Scenarios
It is important that GenAI models are designed and tested under realistic conditions. To do so requires access to realistic production-like data. This data must be both secure, compliant and devoid of PII data to mitigate any unattended data breaches. It is easy to use Broadcom’s data platform replicating secure, compliant, and realistic lower environments. Once the data is secure, it is effortless to produce a variety of virtual copies of production data across complex environment landscapes and distributed data sources. Engineers are able to simulate real-world scenarios, providing targeted insights into its accuracy and performance.
-
Facilitating Multiple Variations
GenAI models need to be designed and tested across a variety of data variations to ensure robustness. DAVE enables teams to effortlessly create multiple virtual copies of their datasets, making it easier to extrapolate variations of your datasets that act as input data streams to the model. By testing with these more diverse variations, teams can evaluate the model’s ability to adapt to different scenarios such as changes in customer behavior, market conditions, or other variables. Organizations are able to focus on what matters and not invest time dealing with the complexities of spinning up servers, storage, databases and the overhead of managing and maintaining environments. DAVE is a centralized way of managing several copies of your data in one single place.
-
Accelerating the Data Pipeline
Traditional methods of creating and managing data can be both slow and manual. Thus leading to data friction, especially when new data needs to be generated for each test cycle every sprint. DAVE significantly reduces the time required to generate relevant test data when you need it while avoiding the cost in time, hardware, and maintenance of spinning up new infrastructure. With the ability to quickly provision environments using virtualized copies of data, development and testing can begin sooner, results can be delivered faster, data environments can be automated, and teams can iterate more quickly. DAVE ensures data is not the bottleneck in your environments.
-
Improving Scalability and Efficiency
As organizations scale, the need for rapid, scalable test data becomes more pressing. In a decentralized environment, different teams may require access to various test data simultaneously. With Broadcom’s DAVE we can help organizations remove bottlenecks, resource contention, environmental complexity, and data friction via secure, scalable, on demand data.
The Road Ahead
Recognizing the growing need for efficient and scalable test data provisioning, Broadcom introduced DAVE to its overall data platform. This solution allows organizations to reduce complexity in provisioning lower environments, removing data friction for engineering, so that organizations can focus on what really matters, product delivery and business outcomes. DAVE plays a pivotal role in enabling decentralized teams to work in parallel without restrictions, reducing data silos, and increasing collaboration. All this is so you can take your processes to the next level as you help your organization to accelerate your GenAI journey.