This project analyzes movie industry data to identify which features are most strongly associated with gross revenue.
The goal of this project is not only to calculate correlations, but to approach the dataset like an analyst: define a clear question, prepare and clean the data, analyze relationships, communicate findings, and explain limitations.
This project follows the Google Data Analytics workflow:
Ask → Prepare → Process → Analyze → Share → Act
Using Python, Pandas, Matplotlib, and Seaborn, I explored a movie industry dataset to understand which factors are most closely related to gross revenue.
The analysis found that budget and votes had the strongest positive relationships with gross revenue. This suggests that movies with larger production budgets and higher audience engagement tend to be associated with higher gross revenue.
Company did not show as strong of a relationship with gross revenue as initially expected. However, because categorical variables were converted into numeric category codes for exploratory correlation analysis, those results should be interpreted carefully.
Which movie features are most strongly associated with gross revenue?
- Do movies with larger budgets tend to generate higher gross revenue?
- Is audience engagement, represented by votes, associated with higher gross revenue?
- Does company appear to have a strong relationship with gross revenue?
- Which numeric movie features have the strongest linear relationship with gross revenue?
- What limitations should be considered when interpreting correlation results?
The dataset used in this project is the Movie Industry dataset from Kaggle.
It includes movie-level information such as:
| Column | Description |
|---|---|
name |
Movie title |
rating |
Movie rating |
genre |
Primary genre |
year |
Original release year column |
released |
Full release date and release country |
score |
Movie score |
votes |
Number of audience/user votes |
director |
Director name |
writer |
Writer name |
star |
Main star |
country |
Country of release/production |
budget |
Movie production budget |
gross |
Gross revenue |
company |
Production/distribution company |
runtime |
Runtime in minutes |
The dataset required cleaning before correlation analysis.
- Checked missing values by column.
- Removed rows where
budgetorgrosswas missing. - Converted
budgetandgrossfrom decimal values to integers. - Created a corrected release year column called
year_correct. - Sorted movies by gross revenue to inspect the highest-grossing films.
- Checked for duplicate rows.
- Reviewed company names for possible spelling or formatting inconsistencies.
Rows with missing budget or gross values were removed because both fields are central to this analysis.
Since the main question focuses on relationships with gross revenue, rows without gross revenue cannot support the analysis. Similarly, rows without budget cannot support the budget-versus-gross comparison.
The dataset includes an original year column, but the released column contains the full release date. To improve reliability, I created a corrected year column by extracting the four-digit year from the released field.
Example:
June 13, 1980 (United States) → 1980
This created a new column:
year_correct
Company names were reviewed for possible inconsistencies. For example, a company could appear in slightly different forms:
Warner Bros.
Warner Brothers
Warner Bros
These values may refer to the same company, but they may also represent different legal entities, divisions, or historical names. For this project, I did not manually merge company names. I only removed exact duplicate rows and documented company-name consistency as a limitation for deeper company-level analysis.
The analysis focused on identifying relationships between movie features and gross revenue.
- Exploratory data analysis with Pandas
- Scatter plot of budget vs gross revenue
- Regression plot using Seaborn
- Pearson correlation analysis
- Correlation heatmaps
- Categorical encoding for exploratory full-feature correlation analysis
This project uses the Pearson correlation coefficient, which measures the strength and direction of a linear relationship between two numeric variables.
Correlation values range from:
| Value | Interpretation |
|---|---|
1.00 |
Strong positive linear relationship |
0.00 |
Little to no linear relationship |
-1.00 |
Strong negative linear relationship |
For this project, the main focus is identifying which features have the strongest relationship with gross.
Categorical variables such as company, genre, rating, and country were converted into numeric category codes so they could be included in the full correlation matrix.
These codes are arbitrary labels and do not represent true numerical rankings. For example, one company code is not “greater than” another company code. Therefore, correlations involving encoded categorical variables should be treated as exploratory rather than definitive.
The notebook includes several visual and statistical outputs:
- Scatter plot: budget vs gross revenue
- Regression plot: budget vs gross revenue
- Numeric correlation matrix
- Numeric correlation heatmap
- Full-feature correlation heatmap using encoded categorical variables
- Ranked correlation results with gross revenue
- Filtered high-correlation pairs
The strongest positive relationships with gross revenue were:
| Feature | Relationship with Gross Revenue |
|---|---|
budget |
Strong positive relationship |
votes |
Moderate positive relationship |
runtime |
Weak positive relationship |
year_correct |
Weak positive relationship |
score |
Weak positive relationship |
The analysis suggests that movies with higher budgets and stronger audience engagement tend to be associated with higher gross revenue.
-
Budget had the strongest positive relationship with gross revenue.
Movies with larger production budgets tended to be associated with higher gross revenue. -
Votes also showed a meaningful positive relationship with gross revenue.
This suggests that audience engagement or popularity may be associated with stronger box office performance. -
Company did not show as strong of a relationship with gross revenue as initially expected.
However, company was converted into numeric category codes, so this result should be interpreted cautiously. -
Correlation does not imply causation.
A strong relationship between budget and gross revenue does not prove that budget alone causes higher revenue. Other factors such as marketing, release timing, franchise popularity, distribution, genre, and audience demand may also influence revenue.
- This analysis measures correlation, not causation.
- Rows with missing
budgetorgrossvalues were removed, which may introduce bias. - Gross revenue was not adjusted for inflation.
- Categorical columns such as
company,genre,rating, andcountrywere converted into numeric category codes for exploratory analysis. These codes are arbitrary and do not represent true rankings. - Company names may contain spelling or formatting inconsistencies.
- The dataset does not include all possible revenue drivers, such as marketing spend, streaming revenue, franchise status, release strategy, or international distribution details.
Future improvements could include:
- Adjusting budget and gross revenue for inflation.
- Creating a
profitcolumn usinggross - budget. - Calculating return on investment using
gross / budget. - Analyzing trends by decade.
- Comparing average gross revenue by genre.
- Building a dashboard in Tableau or Power BI.
- Using more appropriate encoding methods for categorical variables.
- Adding external data such as marketing spend, awards, franchise status, or streaming performance.
- Python
- Pandas
- NumPy
- Matplotlib
- Seaborn
- Jupyter Notebook
- VS Code
- Git
- GitHub
movie-correlation-python-analysis/
├── data/
│ └── movies.csv
├── images/
│ ├── budget_vs_gross_scatter.png
│ ├── budget_vs_gross_regression.png
│ ├── numeric_correlation_heatmap.png
│ └── full_correlation_heatmap.png
├── notebooks/
│ └── movie_correlation_project.ipynb
├── .gitignore
├── requirements.txt
└── README.md
git clone https://github.com/ShayanYawarBhatti/movie-correlation-python-analysis.gitcd movie-correlation-python-analysispython3 -m venv .venvsource .venv/bin/activatepip install -r requirements.txtcode .Then open:
notebooks/movie_correlation_project.ipynb
Run the notebook cells from top to bottom.
This project demonstrates how Python can be used to clean, explore, visualize, and analyze movie industry data.
The analysis found that budget and votes were the strongest positive indicators associated with gross revenue. The project also highlights the importance of careful data preparation, cautious interpretation of correlation results, and clear communication of limitations.
This project was completed as part of my data analytics portfolio to demonstrate structured analytical thinking, Python data analysis, and exploratory data visualization.



