Skip to main content

CST383: Growing My Skills in Python and Data Visualization

This week, I learned more about working with data using Python, Pandas, NumPy, and visualization tools. I already have some experience with coding, so some parts felt familiar, especially reading code, testing outputs, and understanding how variables work. However, this week helped me practice applying those skills specifically to data analysis and visualization.

One important thing I learned was how to choose the correct type of plot based on the variables. For example, a histogram is useful for showing the distribution of one numeric variable, a boxplot is helpful when comparing a numeric variable across categories, and a bar chart or count plot works well for categorical data. I realized that making a graph is not just about writing the code correctly. It is also about understanding what the question is asking and choosing a visualization that clearly answers it.

I also practiced problems involving discrete distributions, such as binomial probability and expected value. These problems helped me understand how probability connects to real situations. For example, the lottery expected value problem showed me that I need to include both the prize and the cost of playing. At first, I was focused only on the winning amount, but after working through it, I understood why the expected winnings can be negative.

Another topic I worked on was campaign contribution data. I practiced grouping data by candidate, occupation, and employment status using functions like groupby(), value_counts(), mean(), median(), and crosstab(). This helped me see how data can be summarized in different ways depending on the question. I also learned that data analysis is not only about getting numbers, but also about explaining what those numbers mean.

A concept I am still working on is deciding the best plot without second guessing myself. I can usually understand the code after seeing it, but I want to get better at choosing the right visualization on my own. I also want to keep practicing stacked bar charts and normalized crosstabs because they are useful, but they require careful interpretation.

Overall, this week helped me build on the coding experience I already have and apply it more toward data science. I learned that before writing code, I should slow down, read the question carefully, identify the variables, decide whether they are numerical or categorical, and then choose the best method. This process will help me become more confident in future data analysis work.

Comments

Popular posts from this blog

CST383: Learning Probability Distributions and Data Visualization in Python

This week, I learned more about probability distributions, density plots, histograms, and how to visualize data using Python libraries such as Pandas, Matplotlib, Seaborn, and SciPy. I practiced creating density plots, box plots, cumulative density plots, and histograms using real datasets. I also learned how changing things like bin width, bandwidth, transparency, and sample size can affect the appearance and interpretation of graphs. Another important topic was understanding skewness and how transformations such as log10 can help make heavily skewed data easier to analyze. One thing I found interesting was how probability density functions (PDFs) and histograms can represent the same data differently. Before this week, I thought graphs mostly showed the same information in different styles, but now I understand that each type of plot has a different purpose and can make patterns easier or harder to notice. I also learned that larger sample sizes tend to reflect the true distribution...

Choosing Between MongoDB and MySQL

This week I learned more about how MongoDB and MySQL are both powerful tools for managing data, but they serve different purposes. MySQL is a relational database that organizes data into tables with rows and columns. It uses SQL (Structured Query Language) to define and manage data, which makes it very structured and reliable. MongoDB, on the other hand, is a NoSQL database that stores data as documents in a flexible JSON-like format . It does not require a fixed schema, so it is easier to change or add new data types as needed. Both databases are similar because they can handle large amounts of data, support indexing for faster searches, and allow users to perform queries to get specific information. They are also widely used in modern applications and can be connected to programming languages like Java, Python, or C++. However, the key difference is how they store and organize data. MySQL is best when data has clear relationships, such as in school systems, banking, or employee ...

CST383 Week 1: Python for Data Science

This week, I learned the basics of Python for data science and how tools like NumPy are used. I already have programming experience from my computer science classes, but Python feels different from languages like Java or C++. It is easier to write and more flexible because it does not require strict data types. This makes coding faster, but I also need to be careful to avoid mistakes. We also learned about the Python data science ecosystem, such as NumPy, Pandas, and tools like Google Colab and Jupyter Notebook. I liked using Google Colab because it is simple and runs in the browser, so I don’t need to install anything. However, I am curious when it is better to use local tools like Spyder or Jupyter instead of Colab. The most important concept for me this week was NumPy. I learned that NumPy arrays are much faster than Python lists because they store data in a continuous block of memory and use the same data type. This connects to what I learned in my algorithms class, where performan...