EDA and Data Visualization December 2026

Q.1: A retail company has customer data containing categorical variables such as city, payment method, and product category, and numerical variables such as purchase amount and customer age. It also has an ordinal variable, customer satisfaction level (Low, Medium, High). The company wants to use this data for machine learning to predict high-value customers. Question: a) Using your knowledge of Label Encoding and One-Hot Encoding: b) Recommend the appropriate encoding method for city, payment method, product category, and customer satisfaction level. Justify why you selected each encoding method. Explain the possible c) advantages and disadvantages (trade-offs) of your choices for machine learning models.

Answer:

Introduction:

There are various types of variables in machine learning, and each type should be converted to a suitable numerical form for effective usage by a machine learning algorithm. The given data about the retail company contains categorical (city, payment method, product category), numerical (purchase amount, customer age), and ordinal (customer satisfaction level) variables. The goal is to achieve a predictive model that can predict high-value customers based on the presented data. Label encoding and one hot encoding are two methods of changing categorical data into numeric data. It is important to distinguish between the ordinal and categorical data and the number of possible categories when selecting the appropriate method to avoid incorrect interpretation. Thus, each data type should be converted using the suitable approach.

 

This is partially solved sample answer

Buy complete NMIMS solved assignments for the December 2026 session.

General/Generic Assignment at just ₹180 per assignment.

Customized/ Unique Assignment at just ₹500 per assignment.

Contact No: +91 9741410271 (WhatsApp)

OR

Mail to: smu.assignment@gmail.com

 

Q.2 (A): A financial company uses real-time data for fraud detection and regulatory reporting. Currently, data quality is checked once every three months, but recent problems have occurred because some data is outdated, inconsistent, or incomplete. The company is considering moving to continuous data quality assessment, but management is concerned about cost and resources. Question: Evaluate the benefits and challenges of moving from periodic to continuous data quality assessment in a financial organization. Justify whether the company should adopt continuous assessment, considering cost, technology requirements, regulatory risks, and improvements in business performance.

Answer:

Introduction:

With regard to a financial company, data quality is essential as it facilitates the performance of tasks related to fraud detection, risk management, customer transactions, and regulatory reporting. In the given company, the practice of assessing the quality of data every three months is not sufficient to avoid mistakes and errors. Such an obsolete, wrong and incomplete information will have a negative impact on processes such as fraud detection, reporting, and financial loss. Continuous evaluation of the quality of data ensures that they remain under control. For that reason, the company must consider adopting a continuous data-quality assessment practice despite its cost and technological demands.

 

Q.2 (B): A company wants to create an interactive sales dashboard showing monthly sales for different regions and product categories. The current dashboard has too many filters, 3D charts, and complicated graphs, making it difficult for senior users to understand. Analyze the problem of too much interactivity and complexity in the dashboard. How can the dashboard be redesigned to reduce cognitive overload while still allowing users to explore detailed sales data?

Answer:

Introduction:

A sales dashboard should be designed to help managers obtain the relevant information at a fast pace and make decisions based on it. However, on the other hand, having too much interaction and graphics in one dashboard proves to be unproductive. In the case of the provided company, the overuse of filters, 3D visualizations and complex graphics contribute to the overall cognitive overload. Senior level users, who need to obtain the key trends, comparisons, and performance indicators tend to get lost in the abundance of information. It is recommended to rethink and redesign the dashboard with an emphasis on visualization rather than interaction. On the other hand, information is supposed to be accessible, although its structure should be different and allow deep analysis.