Hands-On Data Analysis with Pandas

This is the code repository for my book Hands-On Data Analysis with Pandas, published by Packt on July 26, 2019.

The 1st_edition tag contains all materials as they were at time of publishing the first edition.

IMPORTANT NOTE (April 29, 2021):

This is the code repository for the first edition. For the second edition, use this repository instead.

Book Description

Data analysis has become an essential skill in a variety of domains where knowing how to work with data and extract insights can generate significant value.

Hands-On Data Analysis with Pandas will show you how to analyze your data, get started with machine learning, and work effectively with Python libraries often used for data science, such as pandas, NumPy, matplotlib, seaborn, and scikit-learn. Using real-world datasets, you will learn how to use the powerful pandas library to perform data wrangling to reshape, clean, and aggregate your data. Then, you will learn how to conduct exploratory data analysis by calculating summary statistics and visualizing the data to find patterns. In the concluding chapters, you will explore some applications of anomaly detection, regression, clustering, and classification, using scikit-learn, to make predictions based on past data.

By the end of this book, you will be equipped with the skills you need to use pandas to ensure the veracity of your data, visualize it for effective decision-making, and reliably reproduce analysis across multiple domains.

What You Will Learn

Prerequisite: Basic knowledge of Python or past experience with another language (R, SAS, MATLAB, etc.).

Understand how data analysts and scientists gather and analyze data
Perform data analysis and data wrangling in Python
Combine, group, and aggregate data from multiple sources
Create data visualizations with pandas, matplotlib, and seaborn
Apply machine learning algorithms with sklearn to identify patterns and make predictions
Use Python data science libraries to analyze real-world datasets.
Use pandas to solve several common data representation and analysis problems
Collect data from APIs
Build Python scripts, modules, and packages for reusable analysis code.
Utilize computer science concepts and algorithms to write more efficient code for data analysis
Write and run simulations

Chapter 1, Introduction to Data Analysis, will teach you the fundamentals of data analysis, give you a foundation in statistics, and get your environment set up for working with data in Python and using Jupyter Notebooks.
Chapter 2, Working with Pandas DataFrames, introduces you to the pandas library and shows you the basics of working with DataFrames.
Chapter 3, Data Wrangling with Pandas, discusses the process of data manipulation, shows you how to explore an API to gather data, and guides you through data cleaning and reshaping with pandas.
Chapter 4, Aggregating Pandas DataFrames, teaches you how to query and merge DataFrames, perform complex operations on them, including rolling calculations and aggregations, and how to work effectively with time series data.
Chapter 5, Visualizing Data with Pandas and Matplotlib, shows you how to create your own data visualizations in Python, first using the matplotlib library, and then directly from pandas objects.
Chapter 6, Plotting with Seaborn and Customization Techniques, continues the discussion on data visualization by teaching you how to use the seaborn library for visualizing your long form data and giving you the tools you need to customize your visualizations, making them presentation-ready.
Chapter 7, Financial Analysis: Bitcoin and the Stock Market, walks you through the creation of a Python package for analyzing stocks, building upon everything learned in chapters 1-6 and applying it to a financial application.
Chapter 8, Rule-Based Anomaly Detection, covers simulating data and applying everything learned in chapters 1-6 to catching hackers attempting to authenticate to a website, using rule-based strategies for anomaly detection.
Chapter 9, Getting Started with Machine Learning in Python, introduces you to machine learning and building models using the sklearn library.
Chapter 10, Making Better Predictions: Optimizing Models, shows you strategies for improving the performance of your machine learning models.
Chapter 11, Machine Learning Anomaly Detection, revisits anomaly detection on login attempt data, using machine learning techniques, all while giving you a taste of how the workflow looks in practice.
Chapter 12, The Road Ahead, contains resources for taking your skills to the next level and further avenues for exploration.

Notes on Environment Setup

Environment setup instructions are in the chapter 1 of the text. If you don't have the book, you must install Python 3.6 or 3.7, set up a virtual environment, activate it, and then install the packages listed in requirements.txt. You can then launch JupyterLab and use the ch_01/checking_your_setup.ipynb Jupyter notebook to check your setup. Consult this resource if you have issues with using your virtual environment in Jupyter.

Solutions

Each chapter comes with exercises. The solutions for chapters 1-11 can be found here.

About the Author

Stefanie Molin (@stefmolin) is a software engineer and data scientist at Bloomberg in New York City, where she tackles tough problems in information security, particularly those revolving around data wrangling/visualization, building tools for gathering data, and knowledge sharing. She holds a bachelor’s of science degree in operations research from Columbia University's Fu Foundation School of Engineering and Applied Science with minors in Economics and Entrepreneurship and Innovation, as well as a master’s degree in computer science, with a specialization in machine learning, from Georgia Tech. In her free time, she enjoys traveling the world, inventing new recipes, and learning new languages spoken both among people and computers.

Acknowledgements

Since the book limited the acknowledgements to 450 characters, the full version is here.

Name		Name	Last commit message	Last commit date
Latest commit History 228 Commits
.github		.github
_img		_img
appendix		appendix
ch_01		ch_01
ch_02		ch_02
ch_03		ch_03
ch_04		ch_04
ch_05		ch_05
ch_06		ch_06
ch_07		ch_07
ch_08		ch_08
ch_09		ch_09
ch_10		ch_10
ch_11		ch_11
ch_12		ch_12
solutions		solutions
.gitignore		.gitignore
LICENSE		LICENSE
README.md		README.md
acknowledgements.md		acknowledgements.md
apt.txt		apt.txt
environment.yml		environment.yml
requirements.txt		requirements.txt
runtime.txt		runtime.txt

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Hands-On Data Analysis with Pandas

Book Description

What You Will Learn

Table of Contents

Notes on Environment Setup

Solutions

About the Author

Acknowledgements

About

Releases

Packages

Languages

License

ciarabautista/Hands-On-Data-Analysis-with-Pandas

Folders and files

Latest commit

History

Repository files navigation

Hands-On Data Analysis with Pandas

Book Description

What You Will Learn

Table of Contents

Notes on Environment Setup

Solutions

About the Author

Acknowledgements

About

Resources

License

Stars

Watchers

Forks

Releases

Packages 0

Languages

Packages