Weaviate

RAG Pipeline Quickstart with Weaviate

Approximate time to complete: 5-10 minutes, excluding prerequisites

This quickstart will walk you through creating and scheduling a pipeline that uses a web crawler to ingest data from the Vectorize documentation, creates vector embeddings using an OpenAI embedding model, and writes the vectors to a Weaviate vector database.

Before you begin

Before starting, ensure you have access to the credentials, connection parameters, and API keys as appropriate for the following:

A Vectorize account (Create one free here ↗ )
An OpenAI API Key (How to article)
An Weaviate account (Create one on Weaviate ↗)

Step 1: Create a Weaviate Cluster

Create Your Database

Log in to Weaviate, navigate to Clusters, and click Create cluster.
Select a cluster type. For this quickstart, we'll use "Free."
Enter a cluster name, select the cloud region, then click Create.
Save and securely store your cluster's API key.

Step 2: Create a RAG Pipeline on Vectorize

Create a New RAG Pipeline

Open the Vectorize Application Console ↗
From the dashboard, click on + New RAG Pipeline under the "RAG Pipelines" section.
Enter a name for your pipeline. For example, you can name it quickstart-pipeline.
Click on New Vector DB to create a new vector database integration.
Select Weaviate from the list of vector databases.
Enter the parameters in the form using the Weaviate Parameters table below as a guide, then click Create Weaviate Integration.

Weaviate Parameters

Field	Description	Required
Name	A descriptive name to identify the integration within Vectorize.	Yes
Endpoint	The cluster's endpoint.	Yes
API key	The cluster's admin API key.	Yes

Configure the Weaviate integration in your RAG Pipeline

You can think of the Weaviate integration as having two parts to it. The first is authorization with your Weaviate cluster. This part is re-usable across pipelines and allows you to connect to this same application in different pipelines without providing the credentials every time.

The second part is the configuration that's specific to your RAG Pipeline. This is where you specify the name of the table in your Weaviate database. If the table does not already exist, Vectorize will create it for you.

Enter your collection name to complete configuration of your Weaviate integration for your RAG pipeline.

Create Weaviate Integration

Configure AI Platform

Click on New AI Platform.
Select OpenAI from the AI platform options.
In the OpenAI configuration screen:
- Enter a descriptive name for your OpenAI integration.
- Enter your OpenAI API Key.
Leave the default values for embedding model, chunk size, and chunk overlap for the quickstart.

Add Source Connectors

Click on Add Source Connector.

Web Crawler Source

Choose the type of source connector you'd like to use. In this example, select Web Crawler.

Choose Web Crawler

Configure Web Crawler Integration

Name your web crawler source connector, e.g., vectorize-docs.
Set Seed URL(s) to https://docs.vectorize.io.

Configure Web Crawler

Click Create Web Crawler Integration to proceed.

Configure Web Crawler Pipeline

Accept all the default values for the web crawler pipeline configuration:
- Throttle Wait Between Requests: 500 ms
- Maximum Error Count: 5
- Maximum URLs: 1000
- Maximum Depth: 50
- Reindex Interval: 3600 seconds

Web Crawler Pipeline Configuration

Click Save Configuration.

Verify Source Connector and Schedule Pipeline

Verify that your web crawler connector is visible under Source Connectors.
Click Next: Schedule RAG Pipeline to continue.

Verify Source Connector

Schedule RAG Pipeline

Accept the default schedule configuration
Click Create RAG Pipeline.

Schedule RAG Pipeline

Step 3: Monitor and Test Your Pipeline

Monitor Pipeline Creation and Backfilling

The system will now create, deploy, and backfill the pipeline.
You can monitor the status changes from Creating Pipeline to Deploying Pipeline and Starting Backfilling Process.

Pipeline Creation

Once the initial population is complete, the RAG pipeline will begin crawling the Vectorize docs and writing vectors to your Pinecone index.

Pipeline Backfilling

View RAG Pipeline Status

Once the website crawling is complete, your RAG pipeline will switch to the Listening state, where it will stay until more updates are available.

Pipeline Listening State

Test Your Pipeline in the RAG Sandbox

After your pipeline is running, open the RAG Sandbox for the pipeline by clicking the RAG Sandbox link on the Pipeline Details page, or the magnifying glass icon on the RAG Pipelines page.

Open RAG Sandbox

In the RAG Sandbox, you can ask questions about the data ingested by the web crawler.
Type a question into the input field (e.g., "What are the key features of Vectorize?"), and click Submit.

Ask Questions in Sandbox

The system will return the most relevant chunks of information from your indexed data, along with an LLM response.

This completes the RAG pipeline quickstart. Your RAG pipeline is now set up and ready for use with Weaviate and Vectorize.

RAG Pipeline Quickstart with Weaviate​

Before you begin​

Step 1: Create a Weaviate Cluster​

Create Your Database​

Step 2: Create a RAG Pipeline on Vectorize​

Create a New RAG Pipeline​

Weaviate Parameters​

Configure the Weaviate integration in your RAG Pipeline​

Configure AI Platform​

Add Source Connectors​

Configure Web Crawler Integration​

Configure Web Crawler Pipeline​

Verify Source Connector and Schedule Pipeline​

Schedule RAG Pipeline​

Step 3: Monitor and Test Your Pipeline​

Monitor Pipeline Creation and Backfilling​

View RAG Pipeline Status​

Test Your Pipeline in the RAG Sandbox​

RAG Pipeline Quickstart with Weaviate

Before you begin

Step 1: Create a Weaviate Cluster

Create Your Database

Step 2: Create a RAG Pipeline on Vectorize

Create a New RAG Pipeline

Weaviate Parameters

Configure the Weaviate integration in your RAG Pipeline

Configure AI Platform

Add Source Connectors

Configure Web Crawler Integration

Configure Web Crawler Pipeline

Verify Source Connector and Schedule Pipeline

Schedule RAG Pipeline

Step 3: Monitor and Test Your Pipeline

Monitor Pipeline Creation and Backfilling

View RAG Pipeline Status

Test Your Pipeline in the RAG Sandbox