MEng Thesis: Deep Reinforcement Learning Agents
For my first deep learning research project, I explored the training of classical and deep reinforcement learning agents on neuromorphic hardware with biologically plausible algorithms.
Situation
Reinforcement learning (RL) has surpassed human performance in complex decision-making AI tasks (AlphaGo, AlphaZero, AlphaFold). However, the standard machine learning training algorithm, backpropagation, is not biologically plausible, making it incompatible neuromorphic computing.
Task
For my Master’s, I was tasked with exploring the use of Direct Feedback Alignment (DFA) for training classical and deep reinforcement learning agents, comparing their performance to standard backpropagation-trained agents.

Action
I structured my research in three phases:
- Established classical RL algorithms (Q-learning, SARSA) baselines on custom-built games such as Gridworld and Cliff Walker.
- Engineer lightweight deep RL agents using both backpropagation and DFA, training them on classic control problems from OpenAI Gym such as CartPole, MountainCar, and Acrobot.
- Expand the deep RL agents to use convolutional neural networks (CNNs) as function approximators, training them on complex Atari game environments from OpenAI Gym, such as Breakout and Seaquest

Result
My work demonstrated that neuromorphic direct feedback alignment training performs similarly to backpropagation across all problems tested.
Scars
As a self-taught programmer with a background in aeromechanical engineering, this was my first (very) deep dive into AI research. I made many and learned as many lessons. However, the two key takeaways were:
- Define the jobs to be done early, and reevaluate them often. This lesson came about after spending weeks blocked, not knowing how to proceed with an ambiguous task in an unfamiliar domain, only to realise, with the help of my supervisor, that small, incremental steps are better than giant leaps.
- Test everything, as bugs can crop up in the most unexpected places. I learned this the hard way after spending days training my final deep RL agents, only to find that I was not saving the replay memory buffer, leading to the agents learning almost nothing.
Interdisciplinary Integration
The project gave me exposure to several topics:
- Neuromorphic computing and associated biologically plausible training algorithms.
- Deep learning architectures such as ANNs and CNNs.
- Reinforcement learning algorithms and environments.
Technologies & Skills Used
- TensorFlow
- OpenAI Gym
- Reinforcement Learning (Q-Learning, SARSA)
- Direct Feedback Alignment (DFA)
- Data Analysis