ArangoDB: Fast, Scalable, and Multi-Model Data Storage with Jan Steeman and Jan Stücke

This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some of the topics that you will find here.

Support the show!

04 June 2018

ArangoDB: Fast, Scalable, and Multi-Model Data Storage with Jan Steeman and Jan Stücke - Episode 34 - E34

0:00/0:00

Share on social media:

Description
Transcript
Chapters

Summary

Using a multi-model database in your applications can greatly reduce the amount of infrastructure and complexity required. ArangoDB is a storage engine that supports documents, dey/value, and graph data formats, as well as being fast and scalable. In this episode Jan Steeman and Jan Stücke explain where Arango fits in the crowded database market, how it works under the hood, and how you can start working with it today.

Preamble

Hello and welcome to the Data Engineering Podcast, the show about modern data management
When you’re ready to build your next pipeline you’ll need somewhere to deploy it, so check out Linode. With private networking, shared block storage, node balancers, and a 40Gbit network, all controlled by a brand new API you’ve got everything you need to run a bullet-proof data platform. Go to dataengineeringpodcast.com/linode to get a $20 credit and launch a new server in under a minute.
Go to dataengineeringpodcast.com to subscribe to the show, sign up for the newsletter, read the show notes, and get in touch.
Your host is Tobias Macey and today I’m interviewing Jan Stücke and Jan Steeman about ArangoDB, a multi-model distributed database for graph, document, and key/value storage.

Interview

Introduction
How did you get involved in the area of data management?
Can you give a high level description of what ArangoDB is and the motivation for creating it?
- What is the story behind the name?

How is ArangoDB constructed?
- How does the underlying engine store the data to allow for the different ways of viewing it?

What are some of the benefits of multi-model data storage?
- When does it become problematic?

For users who are accustomed to a relational engine, how do they need to adjust their approach to data modeling when working with Arango?

How does it compare to OrientDB?

What are the options for scaling a running system?
- What are the limitations in terms of network architecture or data volumes?

One of the unique aspects of ArangoDB is the Foxx framework for embedding microservices in the data layer. What benefits does that provide over a three tier architecture?
- What mechanisms do you have in place to prevent data breaches from security vulnerabilities in the Foxx code?
- What are some of the most interesting or surprising uses of this functionality that you have seen?