Carbon-Aware Routing for Edge-Cloud LLM Function Calling
September 15, 2026
A three-tier routing framework uses a k-NN predictor in a semantic-lexical embedding space to distribute function-calling queries based on accuracy, latency, and power. It routes queries to the lowest-emission tier, matching cloud-level accuracy while reducing carbon intensity.
HOW THIS AFFECTS YOU
●
builderYou can reduce inference energy costs and carbon footprints by routing function calls to edge hardware.
●
policyThis provides a technical pathway for meeting sustainability and carbon reporting requirements in AI deployments.