Industries
Real-time data platforms for IoT and connected devices
An IoT data platform has to take in telemetry from millions of devices, process it as it arrives, and keep storage costs under control as the fleet grows. Digitalis designs, builds and runs IoT data platforms on Apache Kafka, Apache Cassandra and ClickHouse, with 24×7 support, on any cloud or on-premises.
IoT data breaks platforms built for ordinary applications
Four characteristics make device data harder than application data:
- Constant write volume. Every device reports continuously, so the platform has to be sized for write throughput rather than query load.
- Bursty traffic. Firmware updates, reconnect storms and peak usage hours create spikes many times the normal rate.
- Time-series at scale. Readings must be stored cheaply for years yet still be queryable for recent windows in milliseconds.
- Cost that grows with the fleet. If cost per device is not designed in, every new device makes the platform less profitable.
Solutions for IoT and connected devices
Time-series data at scale
Ingest and store telemetry from millions of devices in real time. See time-series data at scale
Data foundation for analytics and AI
Turn device data into a trusted source for analytics and AI. See data foundation for analytics and AI
A reference architecture for device telemetry
- Apache Kafka takes in device data, absorbs bursts without losing messages and decouples devices from downstream systems.
- Stream processing, such as Apache Flink, enriches and filters events and raises alerts as they arrive.
- Apache Cassandra stores write-heavy device state and time-series data across regions.
- ClickHouse runs fast aggregations across the whole fleet.
- Kubernetes and our 24×7 Managed Service provide portable deployment, autoscaling and round-the-clock operations.
Want help designing the pipeline end to end? See data engineering.
Design principles for IoT platforms that scale
- Partition data by device. This spreads writes evenly and keeps each device's history together.
- Size Kafka for peak traffic. Retention and partitions should absorb a reconnect storm without dropping data.
- Separate hot and cold data. Keep recent readings on fast storage and compress or tier older history to object storage.
- Set retention per data type. Raw telemetry, aggregates and alerts rarely need the same lifetime.
Frequently asked questions
What is an IoT data platform?
An IoT data platform is the set of systems that collects telemetry from connected devices, processes it in real time, stores it and makes it available to applications and analytics. It usually combines a streaming layer, such as Apache Kafka, with databases built for high write volumes, such as Apache Cassandra or ClickHouse.
Why use Kafka for IoT data?
Kafka buffers device data durably, so bursts of traffic don't overwhelm downstream systems and no messages are lost if a consumer falls behind. It also lets several applications read the same telemetry independently.
Cassandra or ClickHouse for time-series data?
They do different jobs. Cassandra suits high-volume writes and fast lookups by device, such as current state or recent history. ClickHouse suits analytical queries across the whole fleet, such as averages, trends and anomalies. Many IoT platforms use both.
Can the platform run on our own cloud or on-premises?
We are cloud-agnostic. We build on open-source technologies that run on any cloud, including hyperscalers such as AWS, Azure and Google Cloud and European sovereign clouds such as Scaleway and IONOS, or in your own data centres, so the platform is not tied to one provider.
Do you support platforms after they are built?
We provide 24×7 managed services for the platforms we build and for platforms built by others. See our 24×7 Managed Service.