4.0

Well, here I am back for yet another attempt. This time seems a bit different. I have decided to start with clarifying my intention behind learning data science.

What is it that has kept me here on the pool side for so long and yet has prevented me from taking a head-first dive into this subject. Let’s take it slow and try and different approach. Let’s try and build some motivation first.

What is it that brought me here?

What do I intend to achieve with this skill? Let’s say it’s a year from today and I have been able to acquire a decent level of competency in data storytelling, now what? How will I put this skill to use? What kind of stories do I want to tell the world?

I think anything more than these two questions will be overkill.

I am going to first answer the two questions above and then try and use the principles I learnt from the book ultra-learning to build a learning pathway for myself.

The purpose of this blog is also going to expand a bit to accommodate all my learning adventures and not just be a diary for the narrow domain of data.

So, let’s begin with my learning goals.

  1. I would like to learn a new language
  2. Data storytelling skills
  3. All that there is to know about Climate Policy and Action
  4. I would like to learn how to produce music
  5. I would like to learn the piano/keyboard

That’s it. These are huge goals but there is this magical thing called atomisation. James Clear did it to habits, I like to do it to any big task or goal.

So this time with intention, let’s start this journey. This space is my learning journal and I’d like to use it as often as possible. For now the goal for that is 3 times a week.

A lot of goals there. Goals are often thought of as sources of pressure. But what if I think of them as just directions. A way to reduce the burden of constant decision making. Let me call them that from now – directions. The above are my 5 directions? Okay that’s a bit confusing. Maybe let’s stick to goals and perceive them in my mind as directions.

All right! Let’s begin!

3.0 Back again

Back on my feet.

I changed gears a bit and started by learning Python first so that I am able some of the code in the book. Datacamp is good but is incredibly incremental in approach. So things move slowly. I am still trying to understand my learning style. While on one hand I know I like to start from the start I also have made the most progress in situations where I’ve had to wing it and build entire projects during a short period of time. I do however know that I need to revisit material many times before I internalize it completely (which is the level of familiarity that I prefer when learning anything.)

So I’ve decided to do both. First the plan is to finish as much as possible in the Python for Data Science track on Datacamp – I am currently on course#2 Intermediate Python – by April 10 (3 months from my 28th birthday). And then do one proper end to end project with as much support from the online community as possible. Then I’ll get back to learning the skills for another month and then return to projects again. This way I keep learning new things while actually applying myself to actual projects.

The dream is that I am prepared enough to take on actual/ new projects by my birthday next year. That would be so cool.

The ultimate idea is to be able to create systems by which to understand the problems that haunt India. To be able to zoom in and out and view policy from the lens of evidence and data.

I also want to improve my GIS and data viz skills along the way.

I am so excited about this phase of my life.

Let’s do this!

2.0 Day 1/100

None of the earlier pursuits worked out as they should have.

Cut to present, here I am again. Definitely not at the starting line. I took a Data Mining class in my second last quarter and it was super tough to say the least. In the process, I did imbibe a little bit of the vocabulary.

So I am starting with ‘O’Reilly’s Hands-On Machine Learning with Scikit-Learn, Keras and Tensorflow’ today.

I am setting the goal of finishing this book in all respects in 100 days including projects. This means I should be done with this book by Dec 27 2020. This blog is going to be my learning journal. This is where I will record all my progress: successes and failures.

The idea is to follow the book chapter for chapter.

Target#1: Today – Sep 18: Chapter 1: The Machine Learning Landscape (34 Pages)

For now

N

Back again

Here we are. I am back again. It’s a new quarter, spring is here! Reading the last post I realised I need to structure and declutter my learning path.

For learning Python, I am choosing DataCamp. They have a Data Scientist Career Track which has 23 courses/ about 90+ hours of learning. However, in order to stay motivated and in touch with context, I also want to learn more about projects to which these methods can be applied.

So there are three parts to my learning path:

Syntax and programming language – Python in DataCamp

Statistical Methods – Program Eval at University and the Stanford course I mentioned last time

Big Picture – I am a little unsure of how to go about this – maybe go through projects on Kaggle or read papers on application of Data Science in the Social Sciences

Let’s see. The number of choices and resources available to learn the subject are just so many, it is overwhelming. I am using this blog to keep me anchored as I slowly navigate through this fascinating yet intimidating world of data science.

Best
N

This is where I start, this is where I stop

I am finally starting this journey formally. I think the name is a little bit of a misnomer, data did not happen completely by accident to me but it also was not the long term plan. Anyway we are here now and I guess it’s a good time to learn to be okay with everything not being perfect. So here goes, welcome to Data By Accident, my learning journal in which I record my small but firm steps into the world of data science.

The idea of this blog came from a video I watched on Youtube by Giles McMullen aka the ‘Python Programmer’. In his video where he discusses resources available to learn Data Science, he emphasises the role a blog can play in one’s learning journey. I think this made a lot of sense. This is not the first time I heard someone talk about logging progress but it did serve the role of final push.

Without further ado, though I am writing after a really long time and would like to ramble on a little bit about how things are shaping up around me right, I am going to jump right into the next steps.

I definitely need some strengthening of Calculus and Statistics. I feel like Calculus can wait but Statistics is a little more urgent as it’s a part of my current coursework at the university. We’ve recently been introduced to regressions and RCTs. So that’s where I start cause that takes care of my assignment and grade too (which has been terrible so far in this quarter). So, first stop is strengthening my understanding of regressions and studying more about RCTs. This will also motivate to think about the bigger problems that I intend to solve with my data science toolkit.

For statistics I am planning to study from Introduction to Econometrics by Stock and Watson. I am also thinking of reading Introduction to Statistical Analysis and pursuing the accompanying MOOC by Stanford of the same name. Since I have a lot of coursework to do as well. I am going to start by committing 5 hours to this per week. I will try and use my university resources to make the most of experience learning statistics.

I then want to start using my Datacamp membership to start using R more. Right now I would still consider myself a beginner in the language as I have used it only for assignments. And even though some of those have been based on real-world data sets, the exposure has been limited. I also studied a little bit of Python last year as a part of the Data Analyst Nanodegree that I briefly pursued on Udacity.

I intend to use spring break to build an immersive data science experience for myself. The ideal scenario would be to attend a couple of Chi Hack nights and go for a meet up or two. And of course get through as many DataCamp courses as I can in th Data Scientist with R career track. It is essential for me to pursue R first and then start with Python again. Actually come to think about it, it might make sense for me to get started with the Python career track during spring break cause that is what I am eventually going to graduate to. Let’s see. I will also need to learn more Stata at some point.

I am considering using a project based approach to learn these languages. Like doing past assignements in STATA and R. I think I am going to have to spend more time constructing my study path cause there are several variables involved. However for now, statistics it is.

As a part of some further commitment making, I will set up a Kaggle and GitHub page next under the same name so that I can record my coding progress there.

See you then.