Guide to Deploy Collate Binaries On-Premises
This guide will help you start using Collate Docker Images to run the Collate OpenMetadata Application in an on-premises Kubernetes cluster, connecting with Argo Workflows for running ingestion from the Collate OpenMetadata Application itself.Architecture
Collate OpenMetadata requires four components:- Collate Server.
- Database: Collate Server stores the metadata in a relational database. Collate supports MySQL or Postgres.
- MySQL version 8.0.42 or greater.
- Postgres version 17.6 or greater.
- Search Engine: OpenSearch 3.4. ElasticSearch is not supported in Collate BYOC because Collate AI relies on OpenSearch’s vector capabilities for Semantic and Hybrid Search.
- Workflow Orchestration: Collate uses Argo Workflows as the orchestrator for ingestion pipelines.
Sizing Requirements
The following sections cover hardware, software, and per-component sizing recommendations for a production-ready deployment.Hardware Requirements
A Kubernetes Cluster with at least 1 Master Node and 3 Worker Nodes is the required configuration. Each Worker Node should have at least:- 4 vCPUs.
- 16 GiB Memory.
- 128 GiB Storage capacity.
Note: If you want Collate workloads scheduled on dedicated nodes, use Kubernetes taints and tolerations. Collate OpenMetadata supports
tolerations via custom Helm values.Software Requirements
- Collate OpenMetadata supports Kubernetes Cluster version 1.24 or greater.
- Collate Docker Images are available via private AWS Elastic Container Registry (ECR). The Collate Team will share credentials and steps to configure Kubernetes to pull Docker Images from AWS ECR.
- For Argo Workflows, Collate OpenMetadata is currently compatible with application version 3.4+.
Database Sizing and Capacity
Collate recommends configuring Postgres. For 100,000 Data Assets and 1,000 Users:- 8 vCPUs.
- 64 GiB Memory.
- 256 GiB Storage Capacity.
- 3,500 IOPS storage.
Search Client Sizing and Capacity
For 100,000 Data Assets and 1,000 Users:- 8 vCPUs.
- 64 GiB Memory.
- 256 GiB Storage Capacity.
Argo Workflows Ingestion Runners
The recommended resources are 4 vCPUs and 16 GiB of Memory.On-Premises Prerequisites
Complete the following object storage setup before installing Argo Workflows and Collate.Object Storage for Argo Workflows Artifacts
Argo Workflows requires object storage to archive ingestion logs. On-premises deployments can use MinIO as an S3-compatible object store.Deploy MinIO (If You Don’t Have an Existing Object Store)
Install MinIO using its Helm chart:Create Kubernetes Secret for MinIO Credentials
Create a Kubernetes secret with the MinIO access credentials:Setup AWS ECR
Collate will provide the credentials to pull Docker Images from a private registry located in AWS ECR.Install AWS CLI
Follow the AWS CLI installation guide to install AWS CLI on your machine.Configure AWS Credentials
Configure the AWS CLI with the credentials for the ECR profile:Kubernetes Docker Registry Secrets for AWS ECR
Create a Kubernetes secret so the cluster can pull images from the private ECR registry:Note: Replace
<<NAMESPACE_NAME>> with the namespace where you want to deploy Collate OpenMetadata Server. If the namespace does not exist yet, create it with kubectl create namespace <<NAMESPACE_NAME>>.Install Argo Workflows
Install Argo Workflows using its Helm chart, then configure it to work with Collate ingestion.Add Helm Repository
Add the community Argo Workflows Helm repository:Create the Argo Namespace
Create a dedicated namespace for Argo Workflows:Kubernetes Secret for Argo Workflows DB Credentials
Create a Kubernetes secret with the database credentials Argo Workflows will use:Create Custom Helm Values for Argo Workflows
Create a file namedargo-workflows.values.yml:
Note: If you are using an existing S3-compatible store (e.g., Ceph, NetApp StorageGRID) instead of MinIO, update
endpoint, bucket, and the secret reference to match your environment. Set insecure: false and configure TLS if your store uses HTTPS.Deploy Argo Workflows
Collate targets application version 3.7.1 using Helm chart version 0.45.23 (Artifact Hub):Optional: Enable Prometheus Metrics
If you have a Prometheus Application running on your cluster, enable metrics using:Install Collate
With Argo Workflows in place, install the Collate OpenMetadata application into the cluster.Create the Collate Namespace
Create a namespace to host the Collate deployment:Kubernetes Service Account for Ingestion
The Collate OpenMetadata Application communicates with Argo Workflows to dynamically trigger ephemeral pods that run ingestion workloads. Create a dedicated Kubernetes Service Account:Create Long-Lived API Token for the ServiceAccount
Create a long-lived token secret for the service account:Configure Kubernetes Roles for the Service Account
Create a fileom-argo-role.yml:
Install the Collate Helm Chart
Create Kubernetes Secrets for the database connection:Note: If you plan to use the DeltaLake connector, the
ARGO_INGESTION_IMAGE value should be:
118146679784.dkr.ecr.eu-west-1.amazonaws.com/collate-customers-ingestion-eu-west-1:om-1.13.0-cl-1.13.0openmetadata.values.yml:
Optional: Enable Prometheus Metrics
Collate Application exposes Prometheus metrics on port8586. Enable the integration using:
Post Installation/Upgrade Steps
Complete the following step after installing or upgrading Collate.Configure ReIndexing
After installation or upgrade, configure ReIndexing from the Collate UI. For detailed steps, refer to the Reindexing Search guide.Troubleshooting
Pods Stuck in Pending State
Check for resource constraints or missing secrets:Argo Workflows Cannot Connect to Object Storage
Verify the MinIO service is reachable from theargo-workflows namespace:
200 OK. If it fails, check that MinIO is running and the endpoint in argo-workflows.values.yml is correct.