Hajar Doukhou

AI Research Engineer · AI Safety

I started by proving theorems,Now I build systems that learn.

And spend as much time thinking about what they optimize for.

Hajar Doukhou

About

Some systems compute. Others learn. The difference interests me, especially the part we can't fully explain yet.

I work on reinforcement learning because it's the framework where the machine learns through consequence, not from a fixed picture of the world. Right now that means studying what happens when you accelerate policy optimization and whether those mechanisms survive contact with language models.

I start with the question before the method. Understanding why something is hard usually reveals which tools are worth reaching for, and which ones will only disguise the problem.

Rigour without impact is just decoration. I want the work to matter.

3 internships3 competition wins2 papers

Work

Selected problems I have chased end to end.

Showing featured projects. More work lives below the fold.

  • Does momentum help or hurt when the objective keeps moving?

    Accelerated PPO Variants · 2026

    Standard PPO uses a clipped surrogate objective to constrain policy updates, but what happens when you add momentum to that? Acceleration mechanisms that work well in supervised learning interact non-trivially with a non-stationary value function and a trust-region constraint. I am implementing and benchmarking accelerated PPO variants against a clean PPO baseline across MuJoCo environments with ten seeds each, tracking not just reward but gradient norms, entropy and policy ratios to understand the optimization dynamics.

    Outcome: Ongoing.

    ResearchEngineering
  • Can a policy learn from logged data without ever exploring?

    Offline RL for Autonomous Driving · 2025

    Real exploration in autonomous driving is unsafe; a bad action is not just noisy training, it can mean a crash. We built a two-stage offline RL pipeline in the DonkeyCar simulator: a convolutional autoencoder compressed raw visual observations into a 32-dimensional latent space, then a TQC expert with gSDE generated high-quality trajectories that served as the fixed dataset for a Conservative Q-Learning agent. Ablations on the latent dimension, CQL conservative weight and dataset quality revealed the conditions under which the offline policy succeeds and where it fails.

    Outcome: The expert reached consistent lap completion in ~70k steps. The offline CQL agent recovered stable driving behavior from the fixed dataset in minutes rather than hours of online interaction.

    ResearchEngineeringSafety
  • What does it take to make federated learning trustworthy end to end?

    Federated Learning as a Lifecycle · 2025

    Federated learning is usually treated as a training loop with privacy bolted on. We argued that in regulated settings, healthcare especially, privacy, governance and auditability need to be structural, not ornamental. I co-authored a paper formalizing FL as a four-phase lifecycle: readiness, training, validation and consensus, deployment and maintenance. Each phase transition is machine-checkable through executable policy rules and produces auditable decision records. Validated on multi-label chest X-ray pathology classification.

    Outcome: The framework formalized FL into 4 auditable phases, making every transition traceable through policy checks and decision records rather than informal operational assumptions.

    ResearchSafety
  • How do you catch and explain pension anomalies without touching the payroll engine?

    Retirement Pension Anomaly Detection and Causal Diagnosis · 2025

    The constraint was strict: no access to the core payroll engine, only its outputs. From 288k records and 214 features, I built an unsupervised detector using a dense autoencoder with per-feature adaptive Huber loss and multi-method thresholding. To move beyond black-box alerts, I added a Bayesian network with a causal hierarchy to attribute root causes and generate traceable explanations auditors could act on. Deployed through FastAPI with a React dashboard and Dockerized services.

    Outcome: 2.5% reconstruction error on payroll data, enabling real-time flagging and root-cause diagnosis while reducing manual verification workload.

    EngineeringSafety
  • Can countries build shared flood intelligence without centralizing their data?

    Federated Flood Prediction for Climate Adaptation · 2025

    Hydrological data in Africa is fragmented across countries and institutions, and centralizing it is often politically or technically unrealistic. I contributed to a federated setup where each country trains locally on discharge and station data while contributing to a continental model through aggregation. The system covers flood-year prediction, magnitude regression, and season classification, with both task-specific and shared multi-head architectures.

    Outcome: 2nd place, AI for Sustainable Cities and Communities Challenge, GITEX Africa 2025.

    ResearchEngineeringSafety

Recognition

First Place

Morocco High-Speed Train Challenge - Rail Industry Summit Casablanca

2024

A vision for railway maintenance that speaks the language of data, combining physics-informed modeling, optimization and LLMs to move from reactive interventions to dynamic planning.

Second Place

AI for Sustainable Cities and Communities Challenge - GITEX Africa - AI for SDGs

2025

Connected cities, protected communities: a federated flood-risk system that connected sensors, geospatial intelligence and local data to strengthen urban resilience and help African territories learn together.

Third Place

Fintech Hackathon - ENSIAS Fintech Club

2024

FortunAI: a step toward more human-centered finance, combining behavioral insights, spending intelligence and real-time guidance to make financial decision-making more accessible and proactive.

Leadership

IEEE ENSIAS Student Branch

@ieee.ensias

President

Aug 2024 – Aug 2025

Sponsorship Manager

Nov 2023 – Jul 2024

Understanding each person's objective function is the prerequisite to building anything together.

Still thinking about what these systems optimize for.

hajar.doukhou@outlook.com

Hajar Doukhou · 2026