Spark integration and analysis with NoSQL Databases 2 - Cassandra

Spark integration and analysis with NoSQL Databases 2 - Cassandra

In this project, we will look at Cassandra and how it is suited for especially in a hadoop environment, how to integrate it with spark, installation in our lab environment.

Videos

Each project comes with 2-5 hours of micro-videos explaining the solution.

Code & Dataset

Get access to 50+ solved projects with iPython notebooks and datasets.

Project Experience

Add project experience to your Linkedin/Github profiles.

What will you learn

Exploratory look at cassandra
Data modelling in Cassandra
Use cases Cassandra in the enterprise
Spark integration using our dataset
Materialized Views
Comparing Analytical queries of MongoDB and Cassandra
Spark Datasources

Project Description

In the last hackerday, we looked at NoSQL databases and their roles in today's enterprise. We talked about design choices with respect to document-oriented and wide-columnar datbases, and conclude by doing hands-on exploration of MongoDB, its integration with spark and writing analytical queries using the MongDB query structures.
Like we also noted, Spark has a benefit of being very extensible to quite a number of storage platforms beyond hadoop. This means that as spark developers, we can write and read from virtually any popular storage platform while building our data pipeline.
In this hackerday, we will conclude that session by take a look at Cassandra. We will look at what it is suited for especially in a hadoop environment, how to integrate it with spark, installation in our lab environment, modelling the UK MOT vehicle testing dataset that we used on MongoDB in the first part. Once loaded, anyone can at anytime, perform analytical queries on the tables.

Similar Projects

In this NoSQL project, we will use two NoSQL databases(HBase and MongoDB) to store Yelp business attributes and learn how to retrieve this data for processing or query.

In this project, we will evaluate and demonstrate how to handle unstructured data using Spark.

In this project, we will look at running various use cases in the analysis of crime data sets using Apache Spark.

Curriculum For This Mini Project