<img height="1" width="1" style="display:none" src="https://www.facebook.com/tr?id=315693165909440&amp;ev=PageView&amp;noscript=1">


The Big Data Blog

Upsolver Receives 2019 Rising Star and Premium Usability Awards from Finances Online

Jun 12, 2019 12:54:45 PM / by Upsolver Team posted in Awards, Finances Online


Read More

A Data Lake Approach to Event Stream Analytics

Jun 4, 2019 9:14:20 PM / by Ori Rafael posted in Database, Data Lake, Event Streams


Read More

ETL Pipelines for Kafka Data: Choosing the Right Approach

May 30, 2019 4:44:59 PM / by Eran Levy posted in Apache Kafka, Streaming Data, ETL

If you’re working with streaming data in 2019, odds are you’re using Kafka - either in its open-source distribution or as a managed service via Confluent or AWS. The stream processing platform, originally developed at LinkedIn and available under the Apache license, has become pretty much standard issue for event-based data, spanning diverse use cases from sensors to application logs to clicks on online advertisements.

Read More

4 Challenges of Using Databases for Streaming Data (and a Solution)

May 21, 2019 1:59:55 PM / by Eran Levy posted in Big Data, Database, Streaming Data


Read More

Big Data Infrastructure: When to Build, When to Buy

May 14, 2019 4:04:53 PM / by Eran Levy posted in Big Data, Data Architecture, Data Engineering

Every software development team makes build-vs-buy decisions on a regular basis. For most coding problems, someone is offering a packaged or white-label solution. The decision whether to purchase a tool or develop an alternative in-house - to ‘build or buy’ - is typically made ad-hoc based on cost, existing engineering skillsets and organizational culture.

Read More

Kafka vs. RabbitMQ: Architecture, Performance & Use Cases

May 7, 2019 2:42:07 PM / by Eran Levy posted in Data Architecture, Apache Kafka, RabbitMQ


Read More

How to Improve AWS Athena Performance: The Complete Guide

Apr 22, 2019 12:59:57 PM / by Eran Levy


Read More

7 Popular Stream Processing Frameworks Compared

Mar 21, 2019 7:03:50 PM / by Eran Levy

Stream processing is a critical part of the big data stack in data-intensive organizations. Tools like Apache Storm and Samza have been around for years, and are joined by newcomers like Apache Flink and managed services like Amazon Kinesis Streams.

Read More

Cloud Data Lake vs On-Premises Data Lake: What You Need to Know

Mar 18, 2019 12:51:45 PM / by Eran Levy posted in Data Lake, Data Architecture, Cloud

Is it time to move your data lake to the cloud? As with any infrastructural choice, there are advantages and trade-offs to deploying in the cloud vs on-premises, and the decision needs to be made on ad-hoc basis based on considerations such as scale, cost, and available technical resources.

Read More

3 Steps To Reduce Your Elasticsearch Costs By 90 - 99%

Feb 27, 2019 4:47:59 PM / by Eran Levy posted in S3, Elasticsearch, Log Analysis

This article covers best practices for reducing the price tag of Elasticsearch using a data lake approach. Want to learn how to optimize your entire streaming data infrastructure? Check out our technical whitepaper to learn how leading organizations generate value from cloud data lakes. Get the paper now!


Elasticsearch is a fantastic log analysis and search tool, used by everyone from tiny startups to the largest enterprises. It’s a robust solution for many operational use cases as well as for BI and reporting, and performs well at virtually any scale - which is why many developers get used to ‘dumping’ all of their log data into Elasticsearch and storing it there indefinitely.

Read More