Computing-Aware Traffic Steering (CATS) optimizes the assignment of user requests to replicated service instances by jointly considering network and computing conditions.
Although CATS has been widely studied for single-service scenarios, cloud--edge continuum deployments are typically multi-tenant: multiple slices share the same distributed computing substrate while pursuing different objectives.
This paper investigates an agentic, Large Language Model (LLM)-assisted architecture for inter-slice computing-resource allocation in heterogeneous multi-tenant CATS.
Slice-specific numerical optimization remains inside local agents, while a centralized LLM-based orchestrator selects CPU allocations across slices and service sites using structured slice descriptions, benchmark definitions, and historical feedback.
To compare slices with different objective functions, we introduce a benchmark-relative gain methodology that expresses each slice performance as an improvement or degradation with respect to a reference allocation.
We implement a proof of concept using an open-source Qwen LLM and a numerical cloud--edge simulator with two representative slice models: weighted end-to-end delay minimization and min-max service-site utilization balancing.
The proposed approach is evaluated against uniform and demand-proportional reference policies in terms of benchmark-relative gains, decision latency, token usage, and output validity.
The results provide an initial assessment of the opportunities and limitations of LLM-assisted coordination for heterogeneous multi-tenant resource management in the cloud--edge continuum.
Keywords: {Cloud--edge continuum, Computing-Aware Traffic Steering, network slicing, multi-tenant resource allocation, large language models, fairness.}

