Blogs

Hi, my name is, what? My name is, who? My name is, chka-chka, Slim Shady (fan).

The spring of 2020, right before the cusp of 18 months of not being able to stay at what had become, in 7 months, a magical place, to which not even Hogwarts could come close, the R-Land. The LBS stadium at R-Land, I longed for it, for so very long. While I was longing for the track and field and the athletics team at Roorkee, a friend of the students, Prof. Ng, was there to take people’s curiosities to teach us about ideas with big, imminent, consequence, which would pan out in couple years after it, the turns of events in human history, through a journey from linear regression on house prices to teaching a machine to detect digits from photos. Obviously, I was imprinted, to the very core, by it, and since then, my career paths have taken a trajectory I could never have wished to take in my childhood as adulthood was on the horizon, for I didn’t know any better, although I did perfectly know, and loved, the mathematics.

In Spring 2026 at UW-Madison, I took a course on the Mathematical Principles of RL. With a core focus on theoretical derivations, I completed an extensive study of the policy gradient algorithms literature, reading papers on CPI (2002), NPG (2001), TRPO (2015), and PPO (2017).

With these foundations I gained from doing my first full-fledged RL project, say on the fun application of Atari games, it became easier to follow the research built on top of these methods and motivated me to pursue this RL field further. Further is the large language models domain which makes full use of policy optimization, where the next token prediction is the action in the RL literature sense. In the Summer of 2026, I did nothing but read more papers ranging from RLHF(2017) to infinity and beyond. This InstructGPT (2022) paper seems to be a good read, and I found it interesting. The reinforcement learning from human feedback premise is given in this paper. There have been numerous methods advancing research in this domain since, including Group Relative Policy Optimization (GRPO), Direct Policy Optimization (DPO), DAPO, and many more to come.

I first got to read in detail about RL while working on a self-motivated project for learning about Multi-Armed Bandits working at HiLabs in September 2023. Although I did enroll in an RL course through NPTEL in August 2022 at Roorkee, for which I forgot to pay the fees, which had some deadline around September 2022, and had to drop out, hah, but during that time of a month or so, I had read at a surface level about it, which was enough to take me back to it. My faint memory of 2021 tells me that I learned about the term reinforcement learning through Coursera.

Favorite book readings

The Elements of Statistical Learning Hastie, Tibshirani, Friedman
Speech and Language Processing Dan Jurafsky, James Martin
A Probabilistic Theory of Pattern Recognition Luc Devroye, László Györfi, and Gábor Lugosi.

My endearment for Music ever since I was a kid

I am very happy when I am listening to music, and this has led me to create a lot of playlists over at spotify. I am also working on a project to make the playlists better, more inclusive of songs, that you may not have already added to your spotify playlists. My spotify insights dashboard can be found here.

Spotify Wrapped, But Everyday

personal Spotify library intelligence from playlists, liked songs, and artist-level curation patterns

recommendation systemsdata visualizationJaccard similaritySpotify API

I built a static Spotify library intelligence dashboard from playlists, liked songs, artist metadata, and saved-track history. It summarizes playlist coverage, curation gaps, artist affinity, playlist diversity, overlap, and listening-memory patterns without requiring a live Spotify login.

The recommendation layer treats missing liked songs as candidate additions and uses lightweight Jaccard-similarity heuristics over playlist and track evidence to suggest where saved songs might belong. It is intentionally interpretable: the page surfaces the overlap reason instead of hiding the recommendation behind a black box.

Cinephile

N-of-1 contextual bandit-based movie recommendation engine

contextual banditsrecommendation systemsN-of-1 systemexploratory MLFlask

Along with being a stan for music, and also, a stan in the literal sense, as per what people say in the pop culture, I am also somewhat of a stan for movies myself, which I attribute to the good taste of movies my friends introduced me to. Also, on the main note, the project ‘Cinephile’ is a contextual bandit-based movie recommendation engine designed as an N-of-1 system where recommendations are driven exclusively by your individual choices rather than collaborative filtering.

By focusing solely on the movie attributes you are drawn toward while integrating a highly tunable exploratory component, Cinephile ensures recommendations adapt to your evolving taste without trapping you in a repetitive feedback loop.

I am always going to the top user of these apps here, but you can be the second! And suggest me how to make them more compelling!

Man, I see in Fight Club the strongest and smartest men who’ve ever lived. I see all this potential, and I see it squandered. Goddamn it, an entire generation pumping gas, waiting tables; slaves with white collars. Advertising has us chasing cars and clothes, working jobs we hate so we can buy shit we don’t need. We’re the middle children of history, man. No purpose or place. We have no Great War. No Great Depression. Our great war is a spiritual war. Our great depression is our lives. We’ve all been raised on television to believe that one day we’d all be millionaires, and movie gods, and rock stars, but we won’t. And we’re slowly learning that fact. And we’re very, very pissed off.

Annual New Year’s Eve Spot: Hilltop Goa

December 2022
Somewhere near Ozran beach; Dil Chahta Hai; Dec 2022

Until It Sleeps
When not typing: 007, Sopranos
Dearest Author: Malcolm Gladwell
Favorite Restaurant: Mom’s Spaghetti

Bangalore, and working at a startup in Bangalore

As I was leaving Bangalore, I knew I would get the change to go there after at least 6-7 years, and I would not hide the fact that I couldn’t hold myself but shed a few bittersweet teers for it. A big silver lining was there, though, as the next stop was my home, Haryana. I was fortunate to start my career in the industry by working at a startup straight out of graduation, at HiLabs. I was lucky to learn under the supervision of a manager who gave me the confidence in my abilities to work on a myriad of problem statements. I became an expert in Git version control. The core product I worked on aimed to automate the ingestion of Medicaid/Medicare roster documents into databases in a standardized format, enabling data interoperability. I also had the opportunity to take on research tasks to extract information for these rosters and store it in structured formats, which I worked through using named entity recognition methods and information extraction methods.