| Pricing & commercial |
| Starting price (monthly) | $300 | Consumption — per instance-hour and per GB |
| Pricing model | Fixed per tier | Three separate consumption meters |
| Cost transparency | High — flat rate | Low — spend split across three services |
| Vendor lock-in | Monthly or annual, no contract | GCP-native; pipelines do not move |
| Integration scope |
| Sources | 260+ | Data Fusion plugin hub; Datastream covers Oracle, MySQL, Postgres, SQL Server |
| Destinations | Warehouses, databases, SaaS, NoSQL, files, APIs, queues, IoT, email | BigQuery, Cloud Storage, Spanner, Cloud SQL, Pub/Sub |
| ETL capabilities | ETL, ELT, Reverse ETL, wildcard processing | ✓Data Fusion (CDAP) for ETL/ELT — no reverse ETL |
| API management | ✓Full | Partial — Apigee, licensed separately |
| EDI processing | ✓X12, EDIFACT, HL7, FHIR | — |
| On-prem deployment | ✓ | —GCP-hosted only |
| Embeddable | ✓ | — |
| Transformations |
| Visual mapping | ✓drag-and-drop designer with live preview | ✓Data Fusion pipeline studio |
| Scripting languages | SQL, JavaScript, Python, XSLT, shell | Java and Python via Beam, SQL in BigQuery |
| Nested and hierarchical data | JSON, XML, Avro, Parquet — read, write, normalize, flatten by dragging | Partial — passes JSON through, unpacking happens downstream |
| Warehouse pushdown (ELT) | ✓transform before load or in the warehouse, same engine | ✓BigQuery |
| Reusable logic | ✓macros, templates, and 3,900+ prebuilt flow templates | ✓Data Fusion plugins |
| Lookups and enrichment | ✓Lookup Builder for cross-source lookups | — |
| Data validation | ✓validation rules with per-step error handling | — |
| Orchestration & workflow |
| Scheduling | Cron expressions and fixed intervals, with per-schedule parameters | Cloud Scheduler or Cloud Composer, separate services |
| Event-driven triggers | ✓HTTP listeners and webhooks, message queues, file and email events | ✓Pub/Sub and Eventarc |
| Continuous execution | ✓looping schedules for CDC and queue consumers | ✓Dataflow streaming jobs |
| Visual workflow builder | ✓Composer canvas, 200+ flow types | ✓Data Fusion pipeline studio |
| Nested workflows | ✓nested flows with conditional and looped steps | Partial — Cloud Composer (Airflow) for real branching |
| Run external tools | ✓shell and SSH scripts, CLI, JavaScript, Python, SQL, HTTP calls | Partial — Cloud Run and Cloud Functions via Workflows |
| Parallel execution | ✓overlapping schedules run as independent, separately cancellable instances | ✓autoscaled parallel workers |
| Retries and error handling | ✓per-step exception handling with notifications | ✓per-service retry policies |
| Run monitoring | ✓per-schedule status, run history, automatic Flow Findings reports | ✓Cloud Monitoring and Logging |
| CDC & Streaming |
| CDC engine | Debezium-compatible, built-in (no Kafka required) | ✓Datastream — log-based, a separate service |
| Database CDC sources | MySQL, Postgres, SQL Server, Oracle, MongoDB, DB2, others | Oracle, MySQL, PostgreSQL, SQL Server |
| Streaming queues | Kafka, EventHubs, Kinesis, SQS, PubSub, ActiveMQ, RabbitMQ | Pub/Sub |
| IoT brokers | ✓MQTT brokers | —IoT Core was retired in 2023 |
| Real-time replication | Log-based CDC, full, incremental | ✓Datastream into BigQuery or Cloud Storage |
| Change tracking modes | Log-based, trigger-based, timestamp/high-watermark | Log-based (Datastream) |
| Developer experience |
| REST API | ✓full API for flows, connections, schedules, and runs | ✓cloud API |
| CLI | ✓full CLI with built-in SQL | ✓cloud CLI |
| MCP server | ✓built-in — connect Cursor, Claude, or ChatGPT to your instance | — |
| Client libraries | ✓Python, Bash, and PowerShell clients | ✓cloud SDKs in many languages |
| Version control | ✓built-in — automatic history, diff, and revert on every artifact | Partial — infrastructure-as-code, not artifact history |
| Embeddable / white-label | ✓ | — |
| Compliance & security |
| SOC 2 Type 2 | ✓audited; report under NDA, SOC 3 public | ✓plus ISO 27001, PCI, FedRAMP |
| HIPAA | ✓supported with a BAA | ✓covered by the cloud BAA |
| GDPR / DPA | ✓compliant, DPA available | ✓ |
| SSO and MFA | ✓SAML SSO, optional 2FA, JWT stateless auth | ✓cloud IAM and directory integration |
| Role and artifact-level access | ✓six roles plus tag-based scoping of flows, connections, and schedules | ✓fine-grained IAM policies |
| Encryption and data handling | ✓TLS in transit, encrypted at rest, customer-managed PGP, SSH tunnels, IP allowlisting; rows are not persisted by default | ✓platform KMS, customer-managed keys |
| Audit logging | ✓admin actions logged, access logs monitored | ✓cloud-native audit trail |
| Security testing | ✓monthly vulnerability and penetration scans, static analysis blocking every build, periodic third-party audits | ✓continuous, platform-wide |
| Gen AI |
| AI agent | ✓Built-in agent (Simba) — builds and edits flows from chat | Partial — Gemini assists with code and SQL — it does not build and run flows |
| Agent capabilities | Reads metadata, reads/samples data, writes JS & SQL, schedules, deploys, monitors | Code assist, SQL generation |
| Natural-language flow building | ✓‘Vibe-build’ — create flows by describing what you want | Partial — Gemini drafts pipeline code |
| AI-driven mapping | ✓Auto-suggests source-to-destination mappings | — |
| Built-in analytics | ✓Agent runs analysis on flow data and pipeline behavior | ✓BigQuery is the analytics layer |
| Chat across product | ✓Same agent context on every screen | Partial — Gemini Cloud Assist |
| CLI for agent | ✓Full CLI access for run/deploy/monitor/manage | ✓gcloud CLI, not agent-driven |
| Trains on customer data | Never | Per Google Cloud terms |