Research

My research addresses the performance, dependability, and cost of distributed software systems and services. I work on decision problems arising in resource provisioning, optimal design, and QoS/SLA management, combining applied probability, optimization, and AI/ML. The three themes below organise my current work; they are also the themes pursued by my group.

QoS management techniques

I study adaptive techniques, algorithms, and control policies for resource allocation, scheduling, and service-level assurance in dynamic, heterogeneous computing environments.

Research interests:

  • Scheduling and orchestration algorithms (serverless, microservices, agentic AI/MCP)
  • Inference serving (early exits, layer skipping, collaborative inference)
  • Total cost of ownership, including cost of inference
  • QoS optimization in cloud continuum architectures (cloud/edge/IoT)

QoS management publications

Modelling and simulation

I develop predictive frameworks and cost models, both analytical and learned, to assess the performance and reliability of distributed computing systems such as enterprise networks, cloud continuum architectures, and their applications.

Research interests:

  • Stochastic models (Markov chains, queueing theory, multi-layer networks)
  • Deep surrogate models and graph neural networks
  • Simulation and statistical inference for model parameterization

Modelling and simulation publications

Dependability and fault-tolerance

I investigate resilient architectures and fault-tolerance mechanisms for AI/ML deployments and for edge and fog computing systems.

Research interests:

  • Anomaly detection
  • Resilient architectures
  • Fault remediation

Dependability publications

Working with me: I supervise PhD students and post-docs interested in modelling (AI/ML, mathematics, applied probability) and its application to computer systems; see my group’s PhD page for application guidelines.