You need to analyze 60,000,000 images stored in JPEG format, each of which is approximately 25 KB. Because you Hadoop cluster isn't optimized for storing and processing many small files, you decide to do the following actions: 1. Group the individual images into a set of larger files 2. Use the set of larger files as input for a MapReduce job that processes them directly with python using Hadoop streaming. Which data serialization system gives the flexibility to do this?
A) CSV
B) XML
C) HTML
D) Avro
E) SequenceFiles
F) JSON

Question

Quizplus · Accepted Answer

SequenceFiles are a data serialization system that allows for the storage of binary key-value pairs. This flexibility makes them ideal for storing and processing large numbers of small files, as they can be used to group the individual images into a set of larger files. Additionally, SequenceFiles can be used as input for a MapReduce job that processes them directly with python using Hadoop streaming.

You Need to Analyze 60,000,000 Images Stored in JPEG Format