●builderImplementing ROCS-style compute sharing can significantly lower the inference latency and cost of your recommendation engine.
●researcherThe decoupling of request and candidate representations via layer masking provides a new way to scale feature-interaction modules.