Skip to content
Back to journal
01 May 2007
6 min read
Game Development

Teaching RPG enemies to learn from the player — a 2007 neural-network prototype

A retrospective on my 2007 final-year project: a C++ and DirectX action RPG that explored online neural-network learning for adaptive NPC behaviour.
閱讀中文版 →
2007 DirectX RPG prototype showing the player and three learning NPC agents

In 2007, my final-year engineering project asked a question that still shapes my work today: could an enemy learn from the way a player behaves instead of following a completely fixed script?

I built a small 3D action RPG prototype in C++ and DirectX. The player fought a group of slime-like agents inside an arena using physical attacks, spells, healing and target locking. Behind that playable demo was an experimental neural network intended to choose when an NPC should wander, follow or evade.

What I built

The project combined three areas that are often treated separately: game-system architecture, real-time graphics and machine learning. The engine used a Win32 game loop with initialise, update and render stages, an object-oriented character controller, and roughly 50 C++ source files with a similar number of headers.

The rendering work included 2D sprites inside a 3D world, camera-facing billboards, animated texture frames, alpha blending and mouse picking. The prototype also included a skybox, lighting, particle effects, camera occlusion, HP and MP systems, frame-rate monitoring and debug information for each agent.

Combat test in the 2007 AI RPG prototype with several NPC agents
The combat test exposed each agent’s health and current action so I could inspect the AI while the game was running.

A small network for moment-to-moment decisions

I deliberately kept the learning model compact enough to update during play. It used four inputs, one hidden layer with three computational neurons, and three possible behavioural outputs.

  • Inputs: distance to the player, the monster’s health ratio, the number of nearby monsters and its current action.
  • Outputs: wander, follow or evade.
  • Learning: feed-forward evaluation followed by error correction and backpropagation.

The larger idea was imitation learning. A developer or player could generate a training set by playing against the agents. Position, health, movement, obstacles and chosen actions would become examples from which an opponent might learn a response. This was an early attempt to treat player behaviour as game data rather than encode every reaction by hand.

What the experiments actually found

The adaptive system was promising, but it was not reliable enough to call finished. That result became the most useful part of the project.

  • Adding more hidden layers did not improve the agents. The deeper versions became less predictable and required more computation.
  • A smaller network gave more direct mappings and was more practical for an in-game experiment, although reducing it too far also produced poor decisions.
  • Adding momentum to the weight-correction process improved stability.
  • Online training depends heavily on the quality and range of the captured player examples. A small or noisy training set cannot produce convincing behaviour.
  • Debug visibility mattered. Showing an agent’s state, target, health and output made the learning system possible to evaluate inside the running game.

The prototype therefore demonstrated the pipeline and the possibility of learning player-like manoeuvres, while also showing that a neural network is not automatically better than authored behaviour. A game still needs clear inputs, measurable outcomes, good training data and fallbacks that protect the player experience.

Castle environment in the 2007 DirectX RPG prototype
The project remained a complete playable graphics prototype even while the adaptive behaviour was still experimental.

Why this 2007 project still matters to me

Many of the conclusions feel current again. Today we can build larger models and collect more data, but interactive AI still has to run within a game loop, explain its state to a developer and produce behaviour that feels fair. Complexity only helps when it improves the player’s experience.

The report also proposed practical tools around the model: an AI test bench, character and map editors, profiling, editable game data and better ways to collect player attributes during play. Those ideas connect directly to how I now think about AI-assisted development: useful intelligence needs an interface, an iteration loop and a clear creative purpose.

Archive note: this retrospective is based on my 2006–07 Bachelor of Engineering final-year report, AI Role Playing Game Development. The screenshots are from the prototype I produced for the project; the text has been rewritten for today’s journal.