Skip to content

05. Data Ingestion and Source Configuration

FirstHive offers flexible and robust data ingestion capabilities to unify your customer data from multiple sources. This chapter covers supported data sources, integration methods, and step-by-step setup instructions to help you get started with your first data source quickly and accurately.

Supported Data Sources

FirstHive supports over 750 diverse data streams across digital and business systems, enabling comprehensive customer data collection.

Digital Sources

  • Website: Analytics, behavioral tracking, form submissions
  • Mobile Applications: In-app events, user interactions, push notifications
  • Social Media: Facebook, Instagram, Twitter, LinkedIn interactions
  • Email Platforms: Campaign performance, engagement metrics

Business Systems

  • CRM Platforms: Customer profiles, sales interactions, support tickets
  • ERP Systems: Transaction history, product data, inventory information
  • Point-of-Sale: In-store purchases, loyalty program data
  • Customer Care: Support interactions, call center data

Integration Methods

FirstHive offers five primary integration modes based on the capabilities of your source or destination systems—whether real-time or batch processing.

  • ETL/ELT/Scripts: Batch data pulls via custom scripts or tools.
  • Pre-configured Connectors: Direct connectors and partner integrations simplify setup.
  • Web and Mobile SDKs: Capture clickstream and event data from websites and apps.
  • API-Led Integrations: Leverage FirstHive REST APIs or 3rd party OEM APIs (REST/SOAP) for flexible, real-time data exchange.
  • Webhooks: Receive real-time, event-driven data pushes from source systems as they occur.

Note: Connecting a new data source is enabled on the backend by FirstHive’s config team — reach out to your Solution Engineer to get a source connected or to make changes to an existing one.

Getting a Data Source Connected

Because integration setup happens on the backend rather than through a customer-facing UI, your Solution Engineer will work with FirstHive’s config team to enable it on your behalf. This covers three pieces of configuration:

  1. Data mapping into events (ingestion) — source fields are mapped into FirstHive’s standard event structure so incoming data lands correctly as events at the point of ingestion. See Events and Entity Schemas for the literal field-level structure this mapping targets.
  2. Uniconfig setup (unification) — the deterministic identity-resolution parameters used to unify incoming events into a single customer identity are configured for your data and vertical. See Unification and Identity Resolution for how identity priority is decided.
  3. Derivation rules setup (entity derivation) — rules that the derivation engine uses to compute Customer, Transaction, and other entities from the underlying event stream are configured for your model. See Events and Entity Schemas for how entities are derived from events.

Once these three are in place, ingestion runs automatically going forward — in real time or batch, depending on the integration method chosen above. To connect a source or request changes to an existing mapping, uniconfig, or derivation rule set, contact your Solution Engineer.

Data Quality and Normalization

FirstHive automatically enhances incoming data with AI-driven cleansing, normalization, and enrichment, powered by the same Model Marketplace described in AI/ML Capabilities.

  • Automated Data Cleansing: Removes errors and standardizes formats.
  • Duplicate Detection: Uses machine learning to identify and merge duplicates.
  • Data Validation: Applies real-time rules and quality scoring.
  • Enrichment: Supplements data with external sources for deeper insights.