Date of Award

Spring 2026

Abstract

Behavioral Cloning (BC) is an imitation learning approach in which a policy is learned from demonstrations represented as sequences of state-action pairs. Standard BC treats all demonstrations as equally trustworthy; consequently, suboptimal or adversarially corrupted demonstrations can disproportionately influence the learned policy. Robust Maximum Entropy Behavioral Cloning (R-MaxEnt BC) addresses this limitation by learning per-demonstration trust weights M and fitting a maximum-entropy policy that is conditioned on these weights.

In this thesis, I propose OR-MaxEnt BC (Optimal Robust Maximum Entropy Behavioral Cloning), a Convex Mixed-Integer Nonlinear Program (MINLP) with optimality guarantees for selecting which demonstrations to trust given the hyperparameter M . Furthermore, I analyze R-MaxEnt BC, correct key errors in prior formulations, and develop optimization frameworks that improve both numerical stability and interpretability. I show that R-MaxEnt BC reduces to multinomial logistic regression under uniform trust, and I derive vectorized objectives and analytical gradients for efficient implementation. To support continuous-action domains without discretization, I extend MaxEnt-based objectives to continuous actions and formulate inference-time action selection as a convex optimization problem.

Empirically, I compare R-MaxEnt BC and OR-MaxEnt BC to Logistic and Linear regression and Discriminator-Weighted Behavioral Cloning (DWBC) across several OpenAI Gym environments under both clean and poisoned demonstration settings. The results demonstrate robustness gains in multiple regimes, while also highlighting sensitivity to M and failure modes under correlated non-expert behavior. [1]

[1] Source code is available at https://github.com/francescomikulis/irl-maxent

Document Type

Master's Thesis

First Advisor

Marek Petrik

Second Advisor

Momotaz Begum

Third Advisor

Wheeler Ruml

Degree Name

Master of Science

Share

COinS