Data foundations Guide

Real-Time Data Streaming in 2026: When It Pays and Which Platform to Choose

Every data platform vendor now sells real-time, and the pitch is that batch is obsolete. The companies with the best public evidence stream only the decisions that cannot wait and run everything else on a schedule. This guide sets out where streaming earns its cost, what the platforms charge in 2026, what changed when Kafka dropped ZooKeeper and IBM bought Confluent, and how to choose and run a platform without the outages that make the news.

For CTOs, heads of data and platform leads deciding whether to move from batch to streaming, or which streaming platform to run.

Published
Reviewed
Reading time
17 min

The short answer

Streaming pays where a machine makes a decision inside a live transaction or session: fraud checks, pricing, dispatch and in-session recommendations. Most dashboards, reports and daily cycles are cheaper on micro-batch or a schedule. Write down how fresh each decision needs its data, stream only that slice, and price platforms on the bytes that cross availability zones as well as the brokers.

Key takeaways

  • Instacart scores only about 1% of its items in real time and cut the compute for availability scoring by about 80% by scoring the rest less often.1
  • A micro-batch every thirty seconds took one ads pipeline's worst-case delay from about ten minutes to thirty seconds, after per-record streaming broke the consistency of its index.2
  • Apache Kafka 4.0 dropped ZooKeeper in March 2025, Kafka 4.2 made queues production-ready in February 2026, and IBM closed its purchase of Confluent for about $11 billion in March 2026.345
  • On Google's managed Kafka, data moving between zones becomes the largest cost once a cluster runs above 20% utilization; AWS charges $0.02 for each GiB that crosses a zone.67
  • Kafka's broker defaults still set min.insync.replicas and default.replication.factor to 1, and security is optional until someone switches it on.89
  • Landing streams in lakehouse tables no longer needs a self-run pipeline: Snowflake and Databricks both ingest directly with data queryable in about five seconds, by their own figures.1011

Data streaming means moving each event, such as a payment, a click, a sensor reading or a row changed in a database, to the systems that need it within seconds of it happening, instead of collecting events and moving them on a schedule. Most streaming estates are built on a log: Apache Kafka or a service that speaks its protocol, where events are kept in order, can be read by many consumers and can be replayed. A stream processor such as Apache Flink or Spark then filters, joins and aggregates the events as they arrive. The question for a business is rarely whether streaming works. It is which decisions are worth the extra cost and operational care, and which platform carries them most cheaply.

  • ~1%of items Instacart scores in real time, a split that cut the compute for availability scoring by about 80%1
  • $0.02what AWS charges for each GiB of data that crosses an availability zone, the line that dominates busy Kafka bills7
  • $11Bthe enterprise value at which IBM bought Confluent, closed on March 17, 20265

Where streaming earns its cost

The documented returns come from decisions that cannot wait for the next batch because money or a customer is already in motion. Instacart moved fraud features from batch to real time so it could act before a loss occurs, which it says "directly reduces millions of fraud-related costs annually", and moved item availability scores from hours old to seconds old.12 DoorDash scaled its Kafka and Flink event platform to hundreds of billions of events a day with four nines of delivery, and cut the delay before events reach its Snowflake warehouse from a day to a few minutes.13 Stripe says its Radar fraud models, trained on more than 70 trillion data points, reduce fraud by 32% on average.14

Payments regulation is the most durable driver, because it moves fraud losses onto the payment provider. Since October 9, 2025, euro-area payment providers must verify the payee before any credit transfer, with other member states following by July 9, 2027.15 Payment fraud across the European Economic Area reached €4.2 billion in 2024, and payment users bore about 85% of credit transfer fraud losses.16 In the UK, mandatory reimbursement of scam victims returned £112 million in its first nine months, 88% of the money lost.17 FedNow had 1,725 participating banks and credit unions by the first quarter of 2026.18 None of these rules sets a screening latency in milliseconds; the pressure comes through liability, and a payee check is a synchronous lookup that needs a fast database more than a streaming platform.

Survey figures are best read as sentiment. In Confluent's 2025 survey of 4,175 IT leaders, 44% reported a fivefold return on data streaming.19 Its 2026 edition, with 4,625 respondents, dropped the return figure and led instead with 72% saying weak real-time infrastructure is stalling AI, 66% unsure of the lineage, timeliness or quality of their data, and only 32% with agentic AI in production.20 Both come from a company that sells streaming, and we found no independent survey of streaming returns.

How fresh each decision needs its data

The companies with the best evidence tier their freshness. Instacart scores every item daily, the trending items hourly, and only about 1% of items in real time, and reports that this cut compute for the model by about 80% while improving its online metrics.1 An InfoQ case study from May 2026 describes an ads index pipeline that tried per-record streaming, found it left the index in partial-update states, and settled on Spark micro-batches every thirty seconds. Its worst-case delay fell from about ten minutes to thirty seconds, and the gain came from removing scheduling gaps.2

AI agents are the newest argument for streaming everything, and the evidence supports a narrower claim. A peer-reviewed benchmark found that outdated retrieved information "substantially reduces response accuracy" in retrieval-augmented generation.21 A September 2026 preprint found that the refresh schedule, more than the age of a cache, governs how stale an agent's answers become.22 The practical rule is to refresh each source at the rate it changes.

Who uses the resultCadence that fitsExample
A machine decides inside a payment, ride, order or live sessionStreaming, sub-second to a few secondsFraud features and item availability at Instacart12
A machine acts on grouped state that must stay consistentMicro-batch, about 10 seconds to 5 minutesThe ads index in the InfoQ case2
A warehouse or lakehouse table read by analysts and jobsMinutes, by streaming ingestion or short micro-batchesDoorDash events reaching Snowflake in minutes13
A dashboard, report or daily planning cycleHourly or daily scheduleInstacart's daily and hourly tiers1
Context for AI agentsRefresh matched to how fast each source changesRefresh scheduling in ChurnBench22
Our synthesis of the cases above. The cadence follows the decision, whatever the tool.

What changed in the platform market in 2025 and 2026

The open-source core changed more in eighteen months than in the five years before. Apache Kafka 4.0, released on March 18, 2025, is the first major release to run entirely without ZooKeeper, and it made the new consumer rebalance protocol generally available.3 Kafka 4.2, in February 2026, made queues (share groups) production-ready, so Kafka can now hand work items to competing consumers as well as keep an ordered log.4 Diskless topics, which keep data in object storage to avoid paying for traffic between zones, were accepted as a design direction in KIP-1150,7 and as of October 2026 no Apache Kafka release ships them.23 Buying object-storage Kafka today means buying it from a vendor.

The commercial map consolidated. IBM completed its purchase of Confluent on March 17, 2026, at $31 a share in cash, an enterprise value of about $11 billion, describing Confluent as used by more than 6,500 enterprises including 40% of the Fortune 500.5 Confluent had bought WarpStream, an object-storage Kafka engine, in 2024,24 so that also now sits with IBM. CoreWeave bought Bufstream from Buf on May 8, 2026.25 The independents responded: Redpanda added object-storage Cloud Topics in March 2026, claiming over 90% savings on cross-zone networking fees,26 Aiven moved Kafka to consumption pricing in August 2026,27 and Aiven, Confluent, Redpanda, StreamNative and Ververica formed a Streamhouse Working Group in September 2026.28 Google's low-cost Pub/Sub Lite has been closed to new customers since September 24, 2024.29

What streaming costs

List prices differ by more than ten times for the same byte, and they only compare through a workload model, because each vendor splits fixed fees, compute, throughput and storage differently. These are the published US list rates we read on October 6, 2026.

OptionFixed or computeData in and outStorage per GB-month
Amazon MSK, Standard brokersFrom $0.204 per broker-hour (m7g.large)Replication between brokers not charged; client traffic across zones billed30$0.10
Amazon MSK, Express brokers$0.408 per broker-hour (m7g.large)$0.01 per GB ingested30$0.10
Amazon Kinesis Data Streams, on-demandPer stream-hour charge$0.08 per GB in, $0.04 per GB retrieved31One day included
Google Managed Service for Apache Kafka$0.09 per DCU-hour (one vCPU with 4 GiB of RAM), $0.054 on a 3-year commitment$0.01 per GiB between zones6Varies by tier
Google Pub/SubNone$40 per TiB of throughput32Billed separately
Confluent Cloud Basic, Standard, Enterprise$0.14, $0.75 and $1.75 to $2.25 per eCKU-hour$0.05; $0.035 to $0.05; $0.02 to $0.05 per GB33$0.08
Confluent Freight (object storage)$2.25 per eCKU-hour, minimum 2$0.014 to $0.03 per GB33$0.03
WarpStream (bring your own cloud)$100, $500 or $1,500 a month by tier$0.01 per GiB written, tiering down to $0.00334Tiered
US list prices read on October 6, 2026. Confluent Dedicated and Redpanda's dedicated tiers do not publish unit rates.

On AWS and Google Cloud the biggest line on a busy cluster is usually data moving between availability zones, more than the brokers themselves. AWS charges $0.02 for each GiB that crosses a zone, Google Cloud charges $0.01 per GiB for replication between zones, and Azure does not charge for it.7 Google's own pricing page states that "for clusters with utilization above 20%, the cost of data transfer between zones is the largest component of the total cost".6 By our arithmetic from AWS list rates, a self-run cluster with three replicas spread over three zones and one consumer group pays about $170 a month in cross-zone fees for every megabyte per second produced, or about $200,000 a year at 100 MB/s before any broker or disk. Amazon MSK does not charge for replication between its brokers, which removes the largest share of that.30

Google's worked example is the cleanest comparison of managed against self-run: at 100 MiB/s, Kafka on Compute Engine comes to about $9,100 a month and the managed service about $11,000.6 The larger savings figures come from vendors selling the fix. WarpStream's founders wrote that inter-zone fees are "70-90% of the cost" of a substantial Kafka workload,35 and Confluent's own 2025 cost comparison lists operations at about $566,000 a month in a column whose total reads about $19,300, so it should not be reused without correction.36 Staffing is the other large cost, and the only published estimates come from vendors: Confluent assumes two full-time engineers for a small production estate and seven to ten at the scale of Lyft's streaming team.37

The cost levers follow from those rates: zone-aware producers and follower fetching so clients read from their own zone, a service that does not bill replication or an object-storage tier for data that can tolerate higher latency, compression, lower replication for data that can be rebuilt such as logs and telemetry, and tiered storage for long retention.

Processing, change data capture and the lakehouse

Apache Flink remains the default open-source engine for stateful, low-latency processing. Flink 2.0, released in March 2025, added a state backend that keeps state in remote storage and, in the project's own tests on heavy queries, ran at 75% to 120% of the throughput of local state; it also removed the DataSet API and dropped Java 8.38 Flink 2.3, in June 2026, added control over where a materialized table resumes when its query changes.39 Managed Flink costs $0.21 per CFU-hour on Confluent Cloud40 and $0.11 per KPU-hour on AWS, plus one extra KPU per application for orchestration.41 Confluent recommends Flink for new work and keeps ksqlDB fully supported for existing applications.42

Spark now competes for low-latency work. Open-source Spark 4.1 shipped a real-time mode for continuous, sub-second processing,43 and Spark 4.2 added real-time support in PySpark and a SQL CHANGES clause for reading change data.44 Databricks reports p99 latencies as low as single-digit milliseconds on its platform, a vendor figure.45 Whatever the engine, the guarantees deserve testing: a peer-reviewed study that injected faults into Kafka Streams, Storm and Flink found duplicate events in the first outputs after Flink restarted, and that joins "often result in unreliable outputs".46

Change data capture, reading changes from a database's own log, is how most operational systems join the stream. Debezium, the main open-source tool, reached version 3.7 in September 2026.47 Its PostgreSQL documentation warns that disk space used by write-ahead log files can spike while a connector falls behind,48 so replication slot lag belongs on the first page of alerts. Managed tools read the same log on a schedule: Fivetran's one-minute syncs are limited to its Enterprise and higher plans.49

If the only consumer is a warehouse or lakehouse, a broker has become optional. Confluent's Tableflow turns Kafka topics into Apache Iceberg tables and became generally available on March 19, 2025,50 and Redpanda's Iceberg Topics did the same in its 25.1 release.51 Snowflake has billed Snowpipe at a flat 0.0037 credits per GB since December 2025,52 and says its Snowpipe Streaming elastic channels make data queryable in as little as five seconds.10 Databricks' Zerobus Ingest, generally available since February 2026, streams into the lakehouse in under five seconds, by Databricks' own figures.11 Real-time analytics on a lakehouse means seconds to minutes.

What breaks in production

Several of Kafka's durability defaults still sit on the unsafe side. In the Kafka 4.3 broker configuration, min.insync.replicas defaults to 1, default.replication.factor to 1 and automatic topic creation to true; unclean leader election, by contrast, defaults to false.8 A producer waiting for every in-sync replica therefore loses nothing only when topics have three replicas and at least two must confirm a write. Exactly-once processing is narrower than it sounds: Flink's documentation warns that if the Kafka transaction timeout is shorter than the longest checkpoint plus restart, "data loss may happen".53 The safer design makes every sink idempotent, keyed on an event ID.

The public incident record points at clients and capacity more than at the log itself. At PagerDuty on August 28, 2025, a new feature created a Kafka producer for every API request, nearly 4.2 million extra producers an hour or 84 times normal; about 95% of events were rejected at the peak, and about 23% of notifications were delayed by five minutes or more over 209 minutes.54 Inngest lost span data in October 2025 when one Kafka cluster's disks filled up.55 An Amazon Kinesis Data Streams event in US-EAST-1 on July 30, 2024 caused errors across AWS services for almost seven hours.56 Cloudflare found that consumers can stay connected and pass health checks while they have stopped making progress, and moved to checks on committed offsets.57

Security has to be switched on. Kafka's documentation states that "security is optional", and it describes encryption in transit, with no native encryption at rest.9 The most serious recent flaw, CVE-2026-33557, let brokers on Kafka 4.1.0 to 4.1.1 configured for OAuth accept any JWT token without validating its signature, issuer or audience; earlier 2025 CVEs abused client login and JAAS settings.58 Monitoring standards are still settling: OpenTelemetry's messaging conventions remain at Development status in version 1.44.0.59 A useful dashboard covers consumer lag in records and seconds, committed-offset progress, rebalance counts, producer counts, broker heap and disk, partitions below their in-sync minimum, and an end-to-end latency probe.

Contracts, schemas and personal data

A stream other teams depend on is an interface and needs the same care as an API. A schema registry rejects incompatible changes; Confluent's defaults to backward compatibility, so consumers upgrade before producers.60 AsyncAPI documents event interfaces the way OpenAPI documents request-response ones, and its 3.1.0 release in January 2026 kept the 3.0 model.61 A data contract adds the owner, quality rules and freshness target for each stream. Confluent's 2026 survey found 66% of IT leaders unsure of the lineage, timeliness or quality of their data,20 which is the problem contracts address.

Personal data in a replayable log is the governance problem with no settled answer. Under UK GDPR an organization has one month to respond to an erasure request, and data that cannot be overwritten at once, such as backups, must be put "beyond use" until it is.62 No regulator has published guidance written for event streams. Retention settings also behave less simply than they look: Datadog found a broker keeping messages well beyond a topic's 36-hour retention because retention applies to closed log segments.63 The defensible design keeps personal data in separate topics with short retention, sends pseudonymized events to long-lived analytics streams, and sets segment size and time so retention means what the policy says.

How to choose a platform

Start by asking whether you need a log at all. A log earns its place when several independent consumers need the same events in order, with replay. For work items that one worker must handle once, a queue is simpler: PostgreSQL's SKIP LOCKED is documented for avoiding lock contention on a queue-like table,64 and Kafka itself now offers queues. To keep a service's database and its events consistent, the standard pattern is a transactional outbox published by change data capture.65

SituationRealistic choice in October 2026
A few MB/s, one cloud, no streaming engineersA serverless or entry managed tier, or Pub/Sub or Kinesis if the stack is native to one cloud
The only consumer is the warehouse or lakehouseDirect ingestion (Snowpipe Streaming, Zerobus), or Tableflow or Iceberg Topics if Kafka already runs
Tens to hundreds of MB/s of logs, telemetry or analytics events on AWS or Google CloudAn object-storage tier (WarpStream, Confluent Freight, Aiven diskless, Redpanda Cloud Topics) or MSK with zone-aware clients
Latency-sensitive transactional streams such as payments and ordersDisk-backed managed Kafka (MSK, Confluent Enterprise or Dedicated, Google Managed Kafka) or Redpanda
On AzureEvent Hubs with its Kafka endpoint, or managed Kafka; zone traffic is free, so the object-storage argument is weaker
Large estate with an existing platform team, or data that must stay on premisesSelf-managed Apache Kafka 4.x, with a support vendor if needed
Concern about vendor ownership and exit termsWeigh IBM (Confluent, WarpStream) and CoreWeave (Bufstream) against the independents and Apache Kafka itself
Our synthesis of the prices, defaults and incidents above. The Kafka protocol, Flink and Iceberg sit under nearly every option, which keeps a later move possible.

Migrations take careful preparation and short downtime windows. Michelin moved a Kafka cluster of 15 brokers, more than 1,500 topics and 5 TB used by more than 25 teams to a managed service, and moved its largest application with under two hours of downtime after extensive preparation. It did not use Confluent's Cluster Linking, which was in preview at the time and did not fit its security constraints.66

Questions to settle before you build or buy

  • Which decisions need data within seconds, which within minutes, and which can wait for a schedule?
  • Does more than one consumer need the same events in order, with replay, or is a queue enough?
  • How many bytes a month will cross availability zones, and who pays for them?
  • Are replication factor, minimum in-sync replicas and topic creation set deliberately on every cluster?
  • Is every sink idempotent, and how will duplicates after a restart be detected?
  • Who owns each stream, what is its schema compatibility rule, and what freshness does it promise?
  • Where does personal data flow, how long is it kept, and how will an erasure request be honored?

The streaming market of 2026 has conceded much of its old pitch. The same vendors that sold real-time as a replacement for batch now sell slower object-storage tiers, tables with a freshness setting and direct ingestion into the lakehouse, because most data does not need milliseconds. The companies that get value stream the decisions that cannot wait, schedule the rest, and run each stream as a product with an owner, a contract and a retention policy.

This is how we approach streaming in our data engineering work: write down a freshness target for each decision, stream only the slice that needs it, price the options on real traffic including the bytes that cross zones, and set safe defaults, idempotent sinks, contracts and alerts before the first consumer goes live.

Questions leaders ask

When should a business use real-time data streaming instead of batch?

When a machine makes a decision inside a live transaction or session, such as a fraud check, a price, a dispatch or an in-session recommendation. Dashboards, reports and daily planning are usually cheaper and simpler on micro-batch or a schedule. Many companies run both, as Instacart does by scoring only about 1% of items in real time.

What does Apache Kafka cost to run in the cloud?

It depends on throughput, retention and where data crosses availability zones. On AWS each GiB crossing a zone costs $0.02, and Google says zone-to-zone transfer is the largest cost above 20% utilization. Google's own example puts 100 MiB/s at about $9,100 a month self-run and $11,000 on its managed service, before staff.

What changed in Kafka in 2025 and 2026?

Kafka 4.0 (March 2025) removed ZooKeeper and made the new consumer rebalance protocol generally available. Kafka 4.2 (February 2026) made queues production-ready. Diskless topics were accepted as a design in KIP-1150 and have not shipped in an Apache release. IBM completed its purchase of Confluent in March 2026.

Is Confluent still independent after the IBM deal?

No. IBM completed the acquisition on March 17, 2026 at $31 a share, about $11 billion in enterprise value. WarpStream, which Confluent bought in 2024, also now belongs to IBM. Apache Kafka itself remains an Apache Software Foundation project.

Do I need Kafka to stream data into Snowflake or Databricks?

Not if the warehouse or lakehouse is the only consumer. Snowpipe Streaming and Databricks Zerobus Ingest accept data directly and make it queryable in seconds by their vendors' figures. Kafka earns its place when several systems need the same events, in order, with replay.

Does Kafka guarantee exactly-once delivery?

Only within limits. Kafka's exactly-once features cover idempotent writes and transactions inside Kafka, and stream processors add their own conditions, such as Flink's warning that a transaction timeout shorter than checkpoint plus restart time can lose data. Design every sink to be idempotent on an event ID.

How long does a move from self-managed Kafka to a managed service take?

The cutover can be short and the preparation long. Michelin moved a cluster with more than 1,500 topics used by more than 25 teams and migrated its largest application with under two hours of downtime, after extensive preparation of topics, schemas, offsets and client switchovers.

Sources

  1. How Instacart modernized the prediction of real-time availability for hundreds of millions of items while saving costsInstacart, 2023
  2. Micro-batch streaming lessons learnedInfoQ, 2026
  3. Apache Kafka 4.0.0 release announcementApache Kafka, 2025
  4. Apache Kafka 4.2.0 release announcementApache Kafka, 2026
  5. IBM completes acquisition of ConfluentIBM Newsroom, 2026
  6. Managed Service for Apache Kafka pricingGoogle Cloud
  7. KIP-1150: Diskless TopicsApache Kafka
  8. Broker configs, Apache Kafka 4.3Apache Kafka documentation
  9. Security overview, Apache Kafka 4.3Apache Kafka documentation
  10. Snowpipe Streaming elastic channels generally availableSnowflake, 2026
  11. Announcing general availability of Zerobus IngestDatabricks, 2026
  12. Lessons learned: the journey to real-time machine learning at InstacartInstacart, 2022
  13. Building scalable real-time event processing with Kafka and Flink at DoorDashInfoQ
  14. Stripe RadarStripe
  15. Instant Payments RegulationEuropean Central Bank
  16. EBA and ECB report on payment fraudEuropean Banking Authority and European Central Bank, 2025
  17. One year on: impact of APP reimbursement on victimsPayment Systems Regulator, 2025
  18. FedNow adoption, economic brief 26-28Federal Reserve Bank of Richmond, 2026
  19. 2025 Data Streaming ReportConfluent, 2025
  20. 2026 Data Streaming ReportConfluent, 2026
  21. HoH: a dynamic benchmark for evaluating the impact of outdated information on retrieval-augmented generationOuyang et al., ACL 2025
  22. ChurnBencharXiv preprint, 2026
  23. KIP-1150 accepted and the road aheadAiven, 2026
  24. Confluent acquires WarpStreamConfluent, 2024
  25. CoreWeave acquires BufstreamBuf, 2026
  26. Redpanda Streaming 26.1 introduces an adaptable streaming engineRedpanda, 2026
  27. Pay for what you stream: Kafka pricingAiven, 2026
  28. Aiven, Confluent, Redpanda, StreamNative and Ververica form Streamhouse Working GroupStreamNative, 2026
  29. Pub/Sub Lite release notesGoogle Cloud
  30. Amazon MSK pricingAmazon Web Services
  31. Amazon Kinesis Data Streams pricingAmazon Web Services
  32. Pub/Sub pricingGoogle Cloud
  33. Confluent Cloud pricingConfluent
  34. WarpStream pricingWarpStream
  35. Kafka is dead, long live KafkaWarpStream, 2023
  36. The true cost of real-time data streamingConfluent, 2025
  37. Understanding and optimizing your Kafka costs, part 2: development and operationsConfluent, 2023
  38. Apache Flink 2.0.0: a new era of real-time data processingApache Flink, 2025
  39. Apache Flink 2.3.0 release announcementApache Flink, 2026
  40. Flink billing on Confluent CloudConfluent
  41. Amazon Managed Service for Apache Flink pricingAmazon Web Services
  42. ksqlDB overviewConfluent documentation
  43. Spark release 4.1.0Apache Spark, 2025
  44. Spark release 4.2.0Apache Spark, 2026
  45. Introducing real-time mode in Apache Spark Structured StreamingDatabricks, 2025
  46. Benchmarking the reliability of stream processing systems under faultsTahir et al., PVLDB 18(3)
  47. Debezium releasesDebezium
  48. Debezium connector for PostgreSQLDebezium documentation
  49. Announcing 1-minute syncsFivetran
  50. Tableflow is now generally availableConfluent, 2025
  51. Redpanda 25.1: Iceberg Topics GARedpanda, 2025
  52. Snowpipe simplified pricingSnowflake release notes, 2025
  53. Apache Kafka connectorApache Flink documentation
  54. August 28 Kafka outages: what happened and how we're improvingPagerDuty, 2025
  55. October incident reportInngest, 2025
  56. Summary of the AWS service event in the Northern Virginia (US-EAST-1) region, July 30, 2024Amazon Web Services
  57. Intelligent, automatic restarts for unhealthy Kafka consumersCloudflare, 2023
  58. Apache Kafka CVE listApache Kafka
  59. Semantic conventions for messaging systemsOpenTelemetry
  60. Schema evolution and compatibilityConfluent documentation
  61. AsyncAPI specification v3.1.0AsyncAPI Initiative, 2026
  62. Right to erasureInformation Commissioner's Office
  63. Lessons learned from running Kafka at DatadogDatadog, 2019
  64. SELECT, PostgreSQL documentationPostgreSQL Global Development Group
  65. Outbox event routerDebezium documentation
  66. Migrate your applications from Kafka on-premises to a managed serviceMichelin IT Engineering, 2022

Written by DigyAi Engineering from the systems we build and run. Every figure links to its public source, and every link and figure was checked on October 6, 2026. No client data appears in our insights.

Read next

All insights
  • A legacy warehouse stands at the back of a plinth with six datasets stacked on its roof. Four are retired before anything moves; the rest run through a reconciliation hall in a governed ring, which shows a green tick, and land as a new drum among open tables on a pine pad. A lit query engine and a second engine both read the same tables, and finally the legacy warehouse's light goes out.

    Data foundations Guide

    Data Warehouse Migration in 2026: When to Move and How to Cut Over Safely

    For CIOs, CDOs and heads of data deciding whether and how to move a Teradata, Oracle, Netezza, SAP BW, Hadoop or Synapse estate to a cloud warehouse or lakehouse.

    16 min read

  • One walled estate of customer data has a gate for each AI use case, and each gate runs the same eight readiness checks: every lamp turns green for a churn model, which is built, while the freshness check turns red for a service agent, whose bar stays down and whose tower is still only drawn.

    Data foundations Checklist

    Is Your Data Ready for AI? A Use-Case Data Readiness Assessment

    For CTOs, chief data officers and CFOs deciding whether the data behind a proposed AI use case can carry it before the budget is committed.

    16 min read

  • Three AI agents send their calls through one lit router, which passes most of them to a fleet of small models and only a hard one to a large frontier model. Each agent has its own budget gauge; one has spent to its cap and a red barrier stops it, while the other two keep working.

    Economics and buying Playbook

    AI Inference Cost: How to Govern LLM and Agent Spend

    For CFOs, CTOs and FinOps leads deciding how to forecast, allocate and cap the recurring cost of LLM applications and AI agents in production.

    16 min read

Get in touch

Tell us what you are building.

Write it as big as you imagine it.