Skip to content
Book Open access

Optimizing Marketing Subsidies via Counterfactual Learning with Asymmetric Reward Function

Jul 2026 · Annual International ACM SIGIR Conference on Research and Development in Information Retrieval · pp. 993-1004 · 2 citations · 71 references
Computer Science

Abstract

In marketing, optimizing subsidy allocation to maximize overall profits is of substantial economic importance. Prior research has employed treatment effect estimation techniques to identify subsidy-sensitive items and design corresponding allocation strategies. However, more accurate treatment effect estimations do not necessarily lead to better allocations, underscoring the critical influence of decision boundaries in decision-making. This paper argues that optimal allocation fundamentally depends on predicting the expected optimal subsidy, a challenge distinct from conventional treatment effect estimation or causal decision-making, which existing approaches fail to address. To fill this gap, we introduce a two-stage Counterfactual optimal subsidy Learning method with an Asymmetric reward (CoLA). In the first stage, we derive a coarse estimate of the expected subsidy threshold by exploiting order information and the conditional independence between expected and observed subsidies. In the second stage, we refine these estimates using an asymmetric loss function, leading to more robust predictions. Under practical budget constraints, we prioritize candidates based on their Sharpe ratios to determine the final subsidy allocation strategy. Experiments on three public datasets and an online A/B test show that our method achieves significant performance improvements, yielding the highest total profit and incremental leverage ratios.

Read PDF

Similar papers

#artificial intelligence Preprint Aug 2026

Learning to Allocate Incentives for Incentivized Advertising via Offline Model-Based Reinforcement Learning

An offline model-based RL framework for cost-controllable sequential incentive allocation is developed and an independent counterfactual scorer evaluates each learned policy on held-out logs, enabling pre-launch selection without costly online exposure.

Zi-Lin Zhao, Han Yang, Tian-Pei Yang et al. · 0 citations
Preprint Sep 2026

Policy Targeting with Market Equilibrium

This paper develops a framework for individualized treatment allocation when interventions shift equilibrium prices and generate spillovers across treated and untreated units. The planner chooses which units receive a subsidy while allowing equilibrium prices to adjust endogenously. We show that the resulting welfare f...

Gyungbae Park · 0 citations
Preprint Sep 2026

Scalable Dynamic Pricing of Substitutable Products through Structure-Guided Policy Learning

Problem definition: We study dynamic pricing of substitutable products with finite, product-specific inventories. Customer substitution couples pricing decisions across products, while the inventory state makes exact dynamic programming intractable at realistic scale. Methodology / results: We develop two MNL-guided po...

Yue Su, Antoine Désir, Axel Parmentier · 0 citations
Preprint Sep 2026

Risk-Averse Welfare Maximization via Marginal Treatment Effects

This paper studies risk-averse treatment allocation when individuals self-select into treatment based on unobserved characteristics. We develop a framework that combines the marginal treatment effect approach to endogenous selection with a general class of coherent risk measures that capture distributional preferences...

Jarrod Burgh, Emerson Melo · 0 citations
Open access Sep 2026

Risk-Adjusted Kelly Investing Under Non-Homogeneous Reward and Risk

In financial applications, the Kelly criterion is a well-known strategy for maximizing long-term growth, but is often criticized for its high-risk approach. To mitigate this effect, previous research proposed two risk-adjusted Kelly criteria in a finite investment horizon: the inflection point, which identifies optimal...

S. Dewasurendra, P. Júdice, Q. Zhu · 0 citations
Open access Aug 2026

Dynamic Learning for Joint Pricing, Advertising, and Inventory Management

Problem definition: Startup firms, often small in size, face the challenge of making cross-functional decisions due to the absence of distinct departments like marketing and operations. These interdependent decisions are further complicated by the lack of historical customer data. As a result, these firms must learn ab...

Huseyin Gurkan, N. Keskin, R. Parker · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.