This is where I start, this is where I stop

I am finally starting this journey formally. I think the name is a little bit of a misnomer, data did not happen completely by accident to me but it also was not the long term plan. Anyway we are here now and I guess it’s a good time to learn to be okay with everything not being perfect. So here goes, welcome to Data By Accident, my learning journal in which I record my small but firm steps into the world of data science.

The idea of this blog came from a video I watched on Youtube by Giles McMullen aka the ‘Python Programmer’. In his video where he discusses resources available to learn Data Science, he emphasises the role a blog can play in one’s learning journey. I think this made a lot of sense. This is not the first time I heard someone talk about logging progress but it did serve the role of final push.

Without further ado, though I am writing after a really long time and would like to ramble on a little bit about how things are shaping up around me right, I am going to jump right into the next steps.

I definitely need some strengthening of Calculus and Statistics. I feel like Calculus can wait but Statistics is a little more urgent as it’s a part of my current coursework at the university. We’ve recently been introduced to regressions and RCTs. So that’s where I start cause that takes care of my assignment and grade too (which has been terrible so far in this quarter). So, first stop is strengthening my understanding of regressions and studying more about RCTs. This will also motivate to think about the bigger problems that I intend to solve with my data science toolkit.

For statistics I am planning to study from Introduction to Econometrics by Stock and Watson. I am also thinking of reading Introduction to Statistical Analysis and pursuing the accompanying MOOC by Stanford of the same name. Since I have a lot of coursework to do as well. I am going to start by committing 5 hours to this per week. I will try and use my university resources to make the most of experience learning statistics.

I then want to start using my Datacamp membership to start using R more. Right now I would still consider myself a beginner in the language as I have used it only for assignments. And even though some of those have been based on real-world data sets, the exposure has been limited. I also studied a little bit of Python last year as a part of the Data Analyst Nanodegree that I briefly pursued on Udacity.

I intend to use spring break to build an immersive data science experience for myself. The ideal scenario would be to attend a couple of Chi Hack nights and go for a meet up or two. And of course get through as many DataCamp courses as I can in th Data Scientist with R career track. It is essential for me to pursue R first and then start with Python again. Actually come to think about it, it might make sense for me to get started with the Python career track during spring break cause that is what I am eventually going to graduate to. Let’s see. I will also need to learn more Stata at some point.

I am considering using a project based approach to learn these languages. Like doing past assignements in STATA and R. I think I am going to have to spend more time constructing my study path cause there are several variables involved. However for now, statistics it is.

As a part of some further commitment making, I will set up a Kaggle and GitHub page next under the same name so that I can record my coding progress there.

See you then.