Anish Acharya, PhD

Research

I am broadly interested in learning theory. My research studies learning as a feedback process under imperfect information. A learning system observes noisy or incomplete feedback, constructs a reward or target, updates its weights, policies, prompts, tools, or coordination structure under constraints, validates the resulting behavior, and changes the observations available in the next round.

I ask what is identifiable from imperfect feedback, how an optimization objective should be constructed when true utility is unobserved, which updates are computationally feasible, and when repeated adaptation is dynamically stable. My current research brings these questions together in models and agent systems that evaluate and modify their own behavior.

01

Recursive Self Improvement

Complete adaptation loop

My long-term goal is to develop learning systems that can adjust their policy (model weights, prompts, memory, tools, and coordination structures) while remaining faithful to their intended objectives. Such systems optimize reward models and evaluators that are necessarily incomplete; as the learner adapts, it can discover and exploit their errors. I study how reward modeling, post-training, constrained search, and validation-based acceptance can be designed together so that repeated adaptation improves true behavior rather than only increasing proxy reward.

Key Questions

  1. How can weights, policies, prompts, tools, and coordination structures be optimized within a common adaptation loop?
  2. How should reward models be learned and updated when true utility is sparse, delayed, or unavailable, and the learner can exploit their errors?
  3. What conditions make repeated updates improve true utility without forgetting, capability collapse, or drift?

Current directions. Reward modeling, post-training and policy optimization, weight and prompt adaptation, orchestration search, and continual self improvement.

02

Distributed & Collaborative Learning

Collective optimization and coordination

Many learning systems consist of participants that hold different data, possess different capabilities, and communicate through a limited or changing topology. My long-term goal is to develop principles for collaborative learning across federated clients, decentralized workers, distributed encoders, autonomous agents, and physical robots. I study how local updates should be aggregated, what information must be communicated, how coordination topology affects convergence, and how collective learning can remain stable under heterogeneity, failures, and adversarial participants.

Key Questions

  1. When can decentralized learners or robot fleets match centralized training under limited communication?
  2. How should aggregation, communication topology, and local computation adapt to heterogeneous or unreliable participants?
  3. Which principles from gossip and federated optimization extend to multi-agent coordination and collaborative policy learning?
03

Robust & Weakly Supervised Learning

Inference from imperfect supervision

This thread studies what can be learned when supervision is corrupted, incomplete, indirect, synthetic, or generated by another model. Robust learning and weak supervision represent different failure modes: in one the signal exists but may be corrupted; in the other the desired signal is only partially observed. My long-term goal is to develop statistically principled and computationally practical methods that distinguish ordinary noise, structured failure, missing information, and valid disagreement across data selection, training, reward modeling, and evaluation.

Key Questions

  1. What remains identifiable under corrupted, incomplete, or model-generated supervision?
  2. Can practical high-dimensional methods attain optimal robustness without manufacturing missing labels?
  3. How should a learner combine supervision sources with different reliability without suppressing valid minority evidence?
04

Efficient Learning

Resource-constrained learning

This thread asks how much data, supervision, communication, memory, computation, and physical interaction are necessary to achieve a desired level of learning performance. I study methods that treat efficiency as part of the learning problem rather than as post-processing: representations learned under memory constraints, distributed coding, sample selection, communication-efficient optimization, and evaluation under fixed model and rollout budgets. The long-term goal is to co-design statistical accuracy and resource allocation across training, inference, and self improvement.

Key Questions

  1. What are the fundamental tradeoffs among statistical accuracy, data, model size, communication, and inference cost?
  2. How can compression, selection, and resource allocation be incorporated into learning rather than applied afterward?
  3. How should self-improving systems allocate tokens, rollouts, judge calls, simulator interactions, and physical trials?
05

Control & Stability

Feedback and stability

This thread supplies the dynamical perspective for the entire program. A continually adapting learner changes the environment, data distribution, and evaluation conditions it will encounter next. Its updates therefore form a closed-loop dynamical system rather than a sequence of independent optimization problems. My long-term interest is in translating ideas from system identification and stability analysis into guarantees for adaptive models and agent systems.

Key Questions

  1. What is the appropriate notion of stability for a learner whose updates change its future data and evaluation conditions?
  2. How can such a system be identified when the learner's interventions alter the environment?
  3. When do locally improving updates produce globally stable behavior rather than drift, oscillation, or collapse?

Full author lists, publication types, and additional links are available under Publications.