{"templateId":"markdown","sharedDataIds":{"sidebar":"sidebar-@l10n/ja/sidebars.yaml"},"props":{"metadata":{"markdoc":{"tagList":["admonition"]},"type":"markdown"},"seo":{"title":"Propensity Scoring (Binary Classifier)","description":"Score every customer with the probability they will convert, churn, or respond, using gradient-boosted models on AI Signals, Treasure AI's ML platform.","siteUrl":"https://docs.treasure.ai","lang":"en-US","jsonLd":{"@context":"https://schema.org","@graph":[{"@type":"Organization","@id":"https://www.treasure.ai/","name":"Treasure AI","url":"https://www.treasure.ai/","logo":"https://www.treasure.ai/hubfs/assets/images/logos/primary-logo.svg"},{"@type":"WebSite","@id":"https://docs.treasure.ai/#website","name":"Treasure AI Documentation","url":"https://docs.treasure.ai/","inLanguage":["en","ja"],"publisher":{"@id":"https://www.treasure.ai/"}}]}},"dynamicMarkdocComponents":[],"compilationErrors":[],"ast":{"$$mdtype":"Tag","name":"article","attributes":{},"children":[{"$$mdtype":"Tag","name":"Heading","attributes":{"level":1,"id":"propensity-scoring-binary-classifier","__idx":0},"children":["Propensity Scoring (Binary Classifier)"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Score every customer on how likely they are to convert, churn, or respond, so campaigns reach the people most likely to act."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"overview-and-use-cases","__idx":1},"children":["Overview and Use Cases"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["The Propensity Scoring solution from Treasure AI's AI Signals platform learns from labeled history and scores every customer with the likelihood of a single outcome you define. You supply a table of customer features and one column marking who did and did not do the thing. It returns a score per customer, a yes/no label, and the diagnostics you need to trust the result."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Propensity Scoring is one of the most flexible options in AI Signals because you choose the outcome. Any question of the form ",{"$$mdtype":"Tag","name":"em","attributes":{},"children":["will this customer do X?"]}," fits, as long as you can label the answer for enough people in the past. Convert, churn, upgrade, open, click, redeem, renew, respond — all use the same model, but with different labels."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Common use cases"]}]},{"$$mdtype":"Tag","name":"ul","attributes":{},"children":[{"$$mdtype":"Tag","name":"li","attributes":{},"children":["Rank prospects by likelihood to convert so sales works the hottest leads first"]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":["Flag customers likely to lapse and trigger proactive win-back journeys"]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":["Target the customers most likely to buy in the next campaign window"]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":["Find customers primed for a premium tier or an add-on"]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":["Predict who is likely to open, click, or unsubscribe, and tune frequency accordingly"]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":["Feed the score into CLTV, NBA, or lookalike targeting as an input signal"]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Who benefits most:"]}," Marketing analysts, CRM managers, and lifecycle teams who have a clear yes/no outcome to predict and want an explainable score across their whole base without building or maintaining a custom ML pipeline."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"how-it-fits-with-the-other-ai-signals","__idx":2},"children":["How It Fits with the Other AI Signals"]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"45%","data-label":"Question you're asking"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Question you're asking"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"55%","data-label":"Use"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Use"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Who matters, based on past purchasing?"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"/ja/products/customer-data-platform/machine-learning/ai-signals/rfm-ai-signals"},"children":["RFM AI Signals"]},": descriptive segments, no modeling required"]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["How likely is this customer to do X?"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Propensity Scoring AI Signals (you are here)"]},": a probability per event you define"]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["How much will this customer be worth?"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"/ja/products/customer-data-platform/machine-learning/ai-signals/cltv-ai-signals"},"children":["CLTV AI Signals"]},": a value forecast and percentile rank"]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Which product should I recommend next?"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"/ja/products/customer-data-platform/machine-learning/ai-signals/nbp-ai-signals"},"children":["NBP AI Signals"]},": a ranked list of products, services, or content"]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Which action should I take for this customer?"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"/ja/products/customer-data-platform/machine-learning/ai-signals/nba-ai-signals"},"children":["NBA AI Signals"]},": a learned per-user policy across candidate actions"]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"em","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Propensity or RFM?"]}," RFM describes what a customer has been, using past purchase behavior to produce segments. That ranking is descriptive: it summarizes history rather than predicting the future, and it cannot tell a loyal customer who is simply between purchase cycles from one who has genuinely lapsed. Propensity Scoring answers the narrower question directly, against an outcome you name."]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"em","attributes":{},"children":["These signals feed each other rather than compete. RFM quartiles and CLTV forecasts make strong input features for a propensity model, and pairing a propensity score with CLTV tells you both how likely a customer is to act and how much that action is worth."]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"quick-start","__idx":3},"children":["Quick Start"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Propensity Scoring runs as two functions on the Treasure AI ML Batch API:"]},{"$$mdtype":"Tag","name":"ul","attributes":{},"children":[{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_train"]}]}," fits a model on your labeled table, tunes it, picks a decision threshold, writes diagnostic tables, and registers the model under ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_name"]},"."]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_predict"]}]}," loads that registered model and scores new customers."]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Your input is a flat table, one row per customer, with a binary target column and any number of feature columns. Column names are yours to choose: point ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["target_column"]}," at your label and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id_column"]}," at your identifier. Everything else in the table is treated as a feature."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["You only need to set a handful of parameters."]}," ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["input_table"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["output_table"]},", and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["target_column"]}," are required for training. Add ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_name"]}," so you can predict with the model later, and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id_column"]}," so scores come back attached to a customer. Everything else has a working default, including ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_type"]},", which defaults to ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["xgboost"]},"."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Feature engineering happens upstream. The solution expects the aggregation and joins to be done already, one row per customer, before the table reaches it."]},{"$$mdtype":"Tag","name":"Admonition","attributes":{"type":"warning","name":"Security"},"children":[{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Store your Treasure AI API key as a Workflow secret named ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["td.apikey"]},", never as a literal value in the workflow definition. Referencing it as a secret keeps the key out of version control, so teams can securely manage sensitive information. See ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"/products/customer-data-platform/data-workbench/workflows/setting-workflow-secrets-from-td-console"},"children":["Setting Workflow Secrets from Treasure Console"]},"."]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"train","__idx":4},"children":["Train"]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"data-language":"yaml","header":{"controls":{"copy":{}}},"source":"# classifier_train.dig\n# Step 1: Train, tune, select a threshold, and register the model under `model_name`\n+train:\n  http>: https://ml-batch-api.treasuredata.com/v1/runs\n  method: POST\n  headers:\n    - authorization: \"TD1 ${secret:td.apikey}\"\n    - X-TD-ML-SESSION-ID: ${session_id}\n    - X-TD-ML-ATTEMPT-ID: ${attempt_id}\n  store_content: true\n  content:\n    input_table: your_database.customer_features\n    output_table: your_database.propensity_output\n    solution_name: classifier_train\n    solution_arguments:\n      model_name: \"lead_scorer_v1\"\n      model_type: \"xgboost\"\n      target_column: \"converted\"\n      user_id_column: \"user_id\"\n      tuning:\n        n_trials: 30\n        optuna_metric: \"fbeta\"\n        fbeta_beta: 0.5\n      threshold:\n        strategy: \"precision_optimal\"\n        min_metric_floor: 0.35\n\n+print_response:\n  echo>: \"Training job submitted. Response: ${http.last_content}\"\n\n# Poll until training completes. The API returns HTTP 408 while the job\n# is running; Treasure Workflow retries until it receives HTTP 200.\n+train_poll_status:\n  http>: https://ml-batch-api.treasuredata.com/v1/runs/${JSON.parse(http.last_content)['id']}/status\n  method: GET\n  headers:\n    - authorization: \"TD1 ${secret:td.apikey}\"\n","lang":"yaml"},"children":[]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Expected output:"]}," a family of diagnostic tables written into the ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["output_table"]}," database, each prefixed ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_"]},". The headline one is ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_run_log"]},", which carries test-set metrics and the selected threshold for the run. The trained model is saved to managed storage under ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_name"]},". See ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"#model-configuration"},"children":["Model Configuration"]}," for the full list and ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"#reading-the-diagnostics"},"children":["Reading the Diagnostics"]}," for how to read them."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"predict","__idx":5},"children":["Predict"]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"data-language":"yaml","header":{"controls":{"copy":{}}},"source":"# classifier_predict.dig\n# Step 2: Predict, load the registered model and write a score per customer\n+predict:\n  http>: https://ml-batch-api.treasuredata.com/v1/runs\n  method: POST\n  headers:\n    - authorization: \"TD1 ${secret:td.apikey}\"\n    - X-TD-ML-SESSION-ID: ${session_id}\n    - X-TD-ML-ATTEMPT-ID: ${attempt_id}\n  store_content: true\n  content:\n    input_table: your_database.customers_to_score\n    output_table: your_database.propensity_scores\n    solution_name: classifier_predict\n    solution_arguments:\n      model_name: \"lead_scorer_v1\"     # Must match the training run\n      user_id_column: \"user_id\"\n\n+print_response:\n  echo>: \"Prediction job submitted. Response: ${http.last_content}\"\n\n+predict_poll_status:\n  http>: https://ml-batch-api.treasuredata.com/v1/runs/${JSON.parse(http.last_content)['id']}/status\n  method: GET\n  headers:\n    - authorization: \"TD1 ${secret:td.apikey}\"\n","lang":"yaml"},"children":[]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Expected output:"]}," one row per customer with a score, a label, the threshold that produced it, and a score bucket for segmentation."]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"header":{"controls":{"copy":{}}},"source":"user_id      score    predicted_label  threshold_applied  score_bucket\n3105285968   0.9412   1                0.41               0.9-1.0\n1850985734   0.6558   1                0.41               0.6-0.7\n274382808    0.2231   0                0.41               0.2-0.3\n358273144    0.0487   0                0.41               0.0-0.1\n"},"children":[]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["There is no separate tuning function. Hyperparameter search happens inside ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_train"]},", so training and prediction are the only two steps."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["To create and schedule these workflows, see ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"/products/customer-data-platform/data-workbench/workflows/getting-started-with-treasure-workflow"},"children":["Getting Started with Treasure Workflow"]},"."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"model-configuration","__idx":6},"children":["Model Configuration"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Each function takes two kinds of input: the data table you point it at, and the parameters you pass it. Both are documented under Data Preparation and Workflow Parameters below."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Training parameters arrive in a core block plus four optional nested blocks — ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["tuning"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["threshold"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["metric"]},", and one model-specific block named after the algorithm. Omit any nested block to accept its defaults. ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"#model-types"},"children":["Model Types"]}," lists each algorithm's complete ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_train"]}," search-range parameters, so you can configure an algorithm from one table."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"classifier_train","__idx":7},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_train"]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_train"]}," fits one model on your labeled table, tunes it, selects a decision threshold, and saves it for reuse."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":4,"id":"data-preparation","__idx":8},"children":["Data Preparation"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["A flat, customer-level table, one row per observation, with a binary target column and any number of feature columns. Column names are yours to choose."]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"18%","data-label":"Field"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Field"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"16%","data-label":"Type"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Type"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"10%","data-label":"Required"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Required"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"56%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["target"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["INTEGER / BOOLEAN"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Yes"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The outcome to learn. Must contain at least two distinct classes; a single-class table fails the run with a user input error. Column name is set via ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["target_column"]},"."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["VARCHAR"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Customer identifier. Excluded from the feature set and carried through to prediction output. Column name is set via ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id_column"]},". Required for distributed prediction."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["feature_1"]}," … ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["feature_N"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["DOUBLE / VARCHAR"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Yes"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Any number of columns describing the customer — demographics, tenure, RFM quartiles, engagement counts, product flags. Every column that is not the target, the identifier, or explicitly dropped is treated as a feature."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["How features are handled"]}]},{"$$mdtype":"Tag","name":"ul","attributes":{},"children":[{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Numeric columns"]}," are used as-is. Missing values pass through to the tree model, which handles them natively; there is no imputation step."]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Categorical (string) columns"]}," are used natively — XGBoost through the pandas ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["category"]}," dtype, CatBoost through ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["cat_features"]}," — with a ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["\"missing\""]}," token standing in for nulls and for values not seen during training."]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Correlated numeric features"]}," are pruned. Pairs above ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["corr_threshold"]}," (default 0.90) have one member dropped, and the decision is recorded in ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_corr_drop_report"]}," alongside correlation and VIF diagnostics."]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["session_id"]}," is dropped automatically."]}," Add anything else you want excluded to ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["columns_to_drop"]},"."]}]},{"$$mdtype":"Tag","name":"Admonition","attributes":{"type":"warning","name":"Data quality requirements"},"children":[{"$$mdtype":"Tag","name":"ul","attributes":{},"children":[{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Provide both classes."]}," The model needs positive and negative examples to learn from. Severe imbalance is handled by ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["class_balance_strategy"]},", but extremely rare positives still produce noisier scores."]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Exclude leakage-prone columns."]}," Drop raw timestamps, identifiers, and any column that is a proxy for the outcome — a ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["converted_date"]}," that only exists for converters, for example. These inflate test metrics and then fail in production. See ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"#preventing-label-leakage"},"children":["Preventing Label Leakage"]},"."]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Do the aggregation upstream."]}," The solution expects one finished row per customer and does no joining or roll-up of its own."]}]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":4,"id":"workflow-parameters","__idx":9},"children":["Workflow Parameters"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Core parameters"]}]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"18%","data-label":"Parameter"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Parameter"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"9%","data-label":"Type"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Type"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"12%","data-label":"Default"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Default"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"9%","data-label":"Required"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Required"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"52%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["input_table"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["—"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Yes"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Feature table to train on, in ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["dbname.table_name"]}," format."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["output_table"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["—"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Yes"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Destination in ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["dbname.table_name"]}," format. Its database is where every ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_*"]}," artifact table is written."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["target_column"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["—"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Yes"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Name of the binary target column. Must have at least two distinct classes."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_name"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["—"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Name the trained model is saved under, unique per Treasure AI account. Not enforced at training time, but ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_predict"]}," cannot load a model that was never named, so set it in practice."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_type"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["xgboost"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["xgboost"]}," or ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["catboost"]},". See ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"#model-types"},"children":["Model Types"]},"."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id_column"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["—"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Identity column, excluded from features. Required if prediction will run across multiple workers."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["columns_to_drop"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["—"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Extra columns to exclude, pipe-separated, for example ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["email|signup_ts"]},". ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["session_id"]}," is already dropped."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["test_size"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["float"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.2"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Fraction held out for honest test metrics. Range 0.05 to 0.5."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["max_load_rows"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["int"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["10000000"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Hard row limit applied when loading training data, as a memory guard. Set 0 to disable. The final model fits on everything loaded."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["tuning"]}," block"]}," — controls the embedded hyperparameter search. Applies to both algorithms."]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"20%","data-label":"Parameter"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Parameter"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"9%","data-label":"Type"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Type"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"11%","data-label":"Default"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Default"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"60%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["n_trials"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["int"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["30"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Number of search trials. Set 0 to skip tuning and use built-in defaults, which is much faster. Raise to 50–100 for a more thorough search."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["cv_folds"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["int"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["3"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Stratified cross-validation folds per trial. Range 2 to 10. More folds are more stable and slower."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["max_tuning_rows"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["int"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["1000000"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Stratified subsample used only during the search. The final model still fits on the full loaded set. Lower it to speed up tuning on very large tables. Set 0 to disable."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["optuna_metric"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["fbeta"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Objective optimized during the search: ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fbeta"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["auc_roc"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["auc_pr"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["precision"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["recall"]},", or ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["log_loss"]},". The aliases ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["auc"]}," and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["roc_auc"]}," map to ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["auc_roc"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["aucpr"]}," and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["pr_auc"]}," map to ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["auc_pr"]},", and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["logloss"]}," maps to ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["log_loss"]},"."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fbeta_beta"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["float"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.5"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Beta for F-beta. Below 1 favors precision, above 1 favors recall. Range 0.01 to 10.0. Also sets the direction of the threshold guardrail."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["early_stopping_rounds"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["int"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Rounds without improvement before a fit stops early, using the evaluation set during cross-validation. 0 turns it off."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["class_balance_strategy"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["balanced"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["balanced"]}," applies automatic positive-class weighting for imbalanced targets. ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["none"]}," leaves the classes as they are."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["threshold"]}," block"]}," — controls how the score-to-label cutoff is chosen."]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"20%","data-label":"Parameter"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Parameter"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"9%","data-label":"Type"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Type"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"18%","data-label":"Default"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Default"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"53%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["strategy"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["precision_optimal"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Operating-point objective: ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["precision_optimal"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["f1_optimal"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["recall_optimal"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["youden_j"]},", or ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fixed"]},". See ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"#threshold-selection"},"children":["Threshold Selection"]},"."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["min_metric_floor"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["float"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.35"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Guardrail on the opposite metric. Acts as a minimum-recall floor when ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fbeta_beta"]}," is below 1, and a minimum-precision floor when it is 1 or above. Range 0.0 to 1.0."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fixed_threshold"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["float"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.5"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The cutoff to use, honored only when ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["strategy"]}," is ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fixed"]},". Range 0.0 to 1.0."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["metric"]}," block"]}," — evaluation and preprocessing controls."]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"20%","data-label":"Parameter"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Parameter"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"9%","data-label":"Type"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Type"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"11%","data-label":"Default"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Default"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"60%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["eval_metric"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["aucpr"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Boosting evaluation metric. Mapped automatically for CatBoost: ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["aucpr"]}," becomes ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["PRAUC"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["auc"]}," becomes ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["AUC"]},", ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["logloss"]}," becomes ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["Logloss"]},"."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["corr_threshold"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["float"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.9"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Pairwise-correlation cutoff. Numeric features above it are dropped and the decision is logged. Range 0.5 to 1.0."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["shap_max_samples"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["int"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["1000"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Maximum rows used to compute SHAP explanations, which bounds the cost of explainability. Minimum 10."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["The remaining block is named after the algorithm — ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["xgboost"]}," or ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["catboost"]}," — and sets the search ranges for that family's hyperparameters. Those, plus the shared boosting ranges, are documented in full under ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"#model-types"},"children":["Model Types"]},"."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":4,"id":"output","__idx":10},"children":["Output"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Training writes a family of diagnostic tables into the ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["output_table"]}," database. Every table is prefixed ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_"]}," and stamped with ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["session_id"]}," and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["created_at"]},", so runs accumulate rather than overwrite each other."]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"34%","data-label":"Table"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Table"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"66%","data-label":"Contents"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Contents"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_run_log"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The headline table. Test-set metrics and run configuration: accuracy, precision, recall, F1, F-beta, AUC-ROC, AUC-PR, the selected threshold, confusion-matrix cells, pseudo-R², chi-square, PSI, the winning hyperparameters, and key config fields."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_train_run_log"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The same schema computed on the training split, so you can compare train against test and spot overfitting."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_feature_list"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The ordered list of feature names the model actually used, with index."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_feature_importance"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["SHAP mean absolute importance per feature, ranked."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_shap_directional"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Per-feature SHAP direction — increases, decreases, or mixed — with mean, median, and standard deviation of SHAP values, the share pushing up and down, and the correlation between feature value and SHAP value."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_roc_curve"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["ROC curve points: false positive rate, true positive rate, threshold."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_prediction_bins"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Histogram of test scores across ten buckets, with counts and percentages."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_vif_scores"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Variance Inflation Factor per numeric feature."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_feature_correlations"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Pairwise feature correlations in long form."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_target_correlations"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Feature-versus-target correlations, ranked."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_corr_drop_report"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Which correlated features were dropped, and why."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_optuna_trials"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The full hyperparameter search log: every trial's parameters and score."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_xgboost_gain_importance"]}," or ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_catboost_gain_importance"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Native gain-based feature importance from the model itself. Which table appears depends on ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_type"]},"."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["The trained model is saved to managed model storage under ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_name"]},". It carries the run configuration, the fitted estimator, the feature names, the winning hyperparameters, the selected threshold, and the preprocessing metadata — category vocabularies and the list of dropped correlated features — so prediction reproduces training exactly."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"classifier_predict","__idx":11},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_predict"]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_predict"]}," loads a registered model and writes a score per customer."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":4,"id":"data-preparation-1","__idx":12},"children":["Data Preparation"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["A customer-level table without the target column, carrying the same feature columns used at training. You do not have to match the training table exactly: missing columns are backfilled, extra columns are dropped, and the remainder is reordered to the trained feature set. Categorical values the model never saw during training map to the ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["\"missing\""]}," token."]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"18%","data-label":"Field"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Field"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"16%","data-label":"Type"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Type"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"10%","data-label":"Required"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Required"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"56%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["VARCHAR"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Customer identifier, carried into the output. Column name is set via ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id_column"]},". Required for distributed prediction; without it the job runs on a single worker and the output falls back to a row index."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["feature_1"]}," … ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["feature_N"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["DOUBLE / VARCHAR"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Yes"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The same features the model was trained on."]}]}]}]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":4,"id":"workflow-parameters-1","__idx":13},"children":["Workflow Parameters"]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"18%","data-label":"Parameter"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Parameter"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"9%","data-label":"Type"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Type"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"12%","data-label":"Default"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Default"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"9%","data-label":"Required"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Required"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"52%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["input_table"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["—"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Yes"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Table of customers to score, in ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["dbname.table_name"]}," format."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["output_table"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["—"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Yes"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Destination for predictions, in ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["dbname.table_name"]}," format."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_name"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["—"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The trained model to load. Its configuration, threshold, and preprocessing all come with it, so there is no ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["target_column"]}," at prediction time. In practice this must match the training run."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id_column"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["—"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Identity column carried into the output. Required for distributed prediction, so that rows shard cleanly and no customer is scored twice."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["columns_to_drop"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["—"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Extra columns to exclude before scoring, pipe-separated."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["output_mode"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["string"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["append"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["No"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["append"]}," or ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["overwrite"]},". Distributed runs are forced to ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["append"]},", since several workers write to the same table."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Prediction is never row-capped. ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["max_load_rows"]}," bounds training only; every row of the prediction input gets scored."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":4,"id":"output-1","__idx":14},"children":["Output"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["One row per scored customer."]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"22%","data-label":"Field"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Field"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"14%","data-label":"Type"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Type"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"64%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["session_id"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["INTEGER"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Workflow session that produced this run. Use it to trace scores back to their job."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_type"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["VARCHAR"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The algorithm that produced the score, ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["xgboost"]}," or ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["catboost"]},"."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["VARCHAR"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Customer identifier from ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id_column"]},". Falls back to a row index if none was given."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["score"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["DOUBLE"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Predicted probability of the positive class, rounded to six decimal places. Higher means more likely."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["predicted_label"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["INTEGER"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["1 when ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["score"]}," is at or above ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["threshold_applied"]},", otherwise 0."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["threshold_applied"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["DOUBLE"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The decision threshold selected during training and carried with the model."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["score_bucket"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["VARCHAR"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Ten-way score bin, for example ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["0.8-0.9"]},". Useful for building segments without writing range logic."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["prediction_timestamp"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["TIMESTAMP"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["When the row was scored."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Using the output."]}," For most marketing work, target by ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["score"]}," rather than ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["predicted_label"]},". Sort descending and take the top decile for a high-value campaign, or suppress the bottom half from expensive paid media. ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["predicted_label"]}," answers a different question — it applies one fixed cutoff chosen during training, which is the right tool when you need a single yes/no decision per customer and the wrong one when you have a fixed budget and want the best N people you can afford. Join the output to your customer table so it becomes a Master Segment attribute, then activate through Journeys, Triggers, Personalization, or Segment Builder."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"model-types","__idx":15},"children":["Model Types"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Propensity Scoring supports two gradient-boosting algorithms behind one interface. They share the same preprocessing, tuning, evaluation, threshold selection, explainability, and output schema, so switching between them is a one-line change and the results stay comparable."]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"16%","data-label":"Algorithm"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Algorithm"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"26%","data-label":"Best for"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Best for"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"32%","data-label":"Strengths"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Strengths"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"26%","data-label":"When to use"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["When to use"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["xgboost"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["General-purpose scoring on rich tabular data."]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["High accuracy, handles non-linear relationships and feature interactions, robust to irrelevant features, scales efficiently to very large tables. Categorical columns are supported natively."]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The default. Start here for most lead-scoring, churn, and conversion problems."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["catboost"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Data with many categorical features."]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Native categorical handling with ordered boosting, which reduces target leakage from high-cardinality categories. Often strong out of the box with less tuning."]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["When your features are heavily categorical — product codes, regions, plan types — or when XGBoost underperforms."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Recommendation:"]}," Start with ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["xgboost"]},". If your feature set is categorical-heavy, train ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["catboost"]}," under a second ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_name"]}," and compare ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_run_log"]}," rows. In both cases, read the SHAP output before activating, to confirm the model is learning sensible signals rather than leaking the answer."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"xgboost--xgboost-","__idx":16},"children":["XGBoost (",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["xgboost"]},")"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["The default. XGBoost builds an ensemble of decision trees, each correcting the errors of the ones before it, using the histogram-based ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["hist"]}," tree method. Categorical columns are handled natively through the pandas ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["category"]}," dtype rather than one-hot expansion, which keeps wide categorical feature sets manageable. Class imbalance is handled by setting ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["scale_pos_weight"]}," automatically when ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["class_balance_strategy"]}," is ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["balanced"]},"."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":4,"id":"model-specific-parameters","__idx":17},"children":["Model-Specific Parameters"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Shared boosting search ranges"]}," — each is a ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["_max"]}," pair inside the ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["xgboost"]}," block."]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"20%","data-label":"Hyperparameter"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Hyperparameter"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"30%","data-label":"Config keys"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Config keys"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"14%","data-label":"Default range"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Default range"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"36%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Boosting rounds"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["n_estimators_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["n_estimators_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["100 – 800"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Number of trees. Widen for more capacity. Minimum 10."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Tree depth"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["max_depth_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["max_depth_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["3 – 10"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Maximum tree depth. Lower to regularize and reduce overfitting. Minimum 1."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Learning rate"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["learning_rate_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["learning_rate_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.01 – 0.3"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Step size, sampled log-uniformly. Lower rates with more rounds usually generalize better."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["XGBoost-specific search ranges"]}]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"20%","data-label":"Hyperparameter"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Hyperparameter"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"32%","data-label":"Config keys"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Config keys"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"14%","data-label":"Default range"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Default range"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"34%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["subsample"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["subsample_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["subsample_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.6 – 1.0"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Row sampling per tree. Lower means more regularization. Valid range 0.1 to 1.0."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["colsample_bytree"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["colsample_bytree_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["colsample_bytree_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.6 – 1.0"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Feature sampling per tree."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["min_child_weight"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["min_child_weight_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["min_child_weight_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["1 – 50"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Minimum sum of instance weight in a child. Higher is more conservative."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["reg_alpha"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["reg_alpha_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["reg_alpha_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.001 – 10.0"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["L1 regularization, sampled log-uniformly."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["reg_lambda"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["reg_lambda_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["reg_lambda_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.001 – 10.0"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["L2 regularization, sampled log-uniformly."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["gamma"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["gamma_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["gamma_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.0 – 10.0"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Minimum split-loss reduction required to make a split."]}]}]}]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":4,"id":"workflow-example","__idx":18},"children":["Workflow Example"]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"data-language":"yaml","header":{"controls":{"copy":{}}},"source":"+train_xgboost:\n  http>: https://ml-batch-api.treasuredata.com/v1/runs\n  method: POST\n  headers:\n    - authorization: \"TD1 ${secret:td.apikey}\"\n    - X-TD-ML-SESSION-ID: ${session_id}\n    - X-TD-ML-ATTEMPT-ID: ${attempt_id}\n  store_content: true\n  content:\n    input_table: your_database.customer_features\n    output_table: your_database.propensity_output\n    solution_name: classifier_train\n    solution_arguments:\n      model_name: \"churn_xgb_v1\"\n      model_type: \"xgboost\"\n      target_column: \"churned\"\n      user_id_column: \"user_id\"\n      columns_to_drop: \"email|signup_ts\"\n      tuning:\n        n_trials: 50\n        optuna_metric: \"auc_pr\"\n        cv_folds: 5\n      threshold:\n        strategy: \"f1_optimal\"\n      metric:\n        corr_threshold: 0.85\n      xgboost:\n        max_depth_min: 3\n        max_depth_max: 8\n        subsample_min: 0.7\n","lang":"yaml"},"children":[]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"catboost--catboost-","__idx":19},"children":["CatBoost (",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["catboost"]},")"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["CatBoost is a gradient-boosting library built around categorical features. It encodes them using ordered target statistics, computed over a random permutation of the data so that a row's own label never contributes to its own encoding. That is what keeps high-cardinality categories — postal codes, SKUs, campaign IDs — from quietly leaking the answer, which is the usual failure mode when those columns are target-encoded by hand. Class imbalance is handled by CatBoost's ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["Balanced"]}," automatic class weights when ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["class_balance_strategy"]}," is ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["balanced"]},"."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":4,"id":"model-specific-parameters-1","__idx":20},"children":["Model-Specific Parameters"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Shared boosting search ranges"]}," — each is a ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["_max"]}," pair inside the ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["catboost"]}," block."]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"20%","data-label":"Hyperparameter"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Hyperparameter"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"30%","data-label":"Config keys"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Config keys"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"14%","data-label":"Default range"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Default range"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"36%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Boosting rounds"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["n_estimators_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["n_estimators_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["100 – 800"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Number of trees, mapped to CatBoost's ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["iterations"]},". Minimum 10."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Tree depth"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["max_depth_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["max_depth_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["3 – 10"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Maximum tree depth. Lower to regularize and reduce overfitting. Minimum 1."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Learning rate"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["learning_rate_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["learning_rate_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.01 – 0.3"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Step size, sampled log-uniformly."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["CatBoost-specific search ranges"]}]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"22%","data-label":"Hyperparameter"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Hyperparameter"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"32%","data-label":"Config keys"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Config keys"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"14%","data-label":"Default range"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Default range"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"32%","data-label":"Description"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Description"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["l2_leaf_reg"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["l2_leaf_reg_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["l2_leaf_reg_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["1.0 – 10.0"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["L2 regularization on leaf values, sampled log-uniformly."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["bagging_temperature"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["bagging_temperature_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["bagging_temperature_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["0.0 – 1.0"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Bayesian bagging intensity. Higher adds more randomness."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["random_strength"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["random_strength_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["random_strength_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["1e-08 – 10.0"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Randomness added to split scoring, sampled log-uniformly."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["border_count"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["border_count_min"]}," / ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["border_count_max"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["32 – 255"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Number of splits considered for numeric features. Minimum 1."]}]}]}]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":4,"id":"workflow-example-1","__idx":21},"children":["Workflow Example"]},{"$$mdtype":"Tag","name":"CodeBlock","attributes":{"data-language":"yaml","header":{"controls":{"copy":{}}},"source":"+train_catboost:\n  http>: https://ml-batch-api.treasuredata.com/v1/runs\n  method: POST\n  headers:\n    - authorization: \"TD1 ${secret:td.apikey}\"\n    - X-TD-ML-SESSION-ID: ${session_id}\n    - X-TD-ML-ATTEMPT-ID: ${attempt_id}\n  store_content: true\n  content:\n    input_table: your_database.customer_features\n    output_table: your_database.propensity_output\n    solution_name: classifier_train\n    solution_arguments:\n      model_name: \"churn_cat_v1\"\n      model_type: \"catboost\"\n      target_column: \"churned\"\n      user_id_column: \"user_id\"\n      tuning:\n        n_trials: 50\n        optuna_metric: \"auc_pr\"\n      threshold:\n        strategy: \"f1_optimal\"\n      catboost:\n        border_count_min: 64\n        border_count_max: 255\n        l2_leaf_reg_min: 2.0\n","lang":"yaml"},"children":[]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"relevant-topics","__idx":22},"children":["Relevant Topics"]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"hyperparameter-tuning","__idx":23},"children":["Hyperparameter Tuning"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Unlike NBA, Propensity Scoring has no separate tuning function. The search runs inside ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_train"]}," every time, so a single job produces a tuned model."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Every trial is written to ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_optuna_trials"]},", so you can see the search rather than just its winner. That table is the place to check whether the search converged or whether the best trial was an outlier."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Three levers control cost against quality:"]},{"$$mdtype":"Tag","name":"ul","attributes":{},"children":[{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["n_trials"]}]}," is the main one. The default of 30 is a reasonable balance. Raise it to 50–100 when model quality matters more than runtime. Set it to 0 to skip the search entirely and train on built-in defaults, which is useful for a fast first look at a new dataset."]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["cv_folds"]}]}," trades stability for time. Three folds is fast; five is steadier on smaller or noisier data."]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["max_tuning_rows"]}]}," caps only the search. Lowering it speeds up tuning on a very large table without changing what the final model is fit on."]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Choose ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["optuna_metric"]}," to match your problem rather than accepting the default reflexively. ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fbeta"]}," with ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fbeta_beta"]}," below 1 biases toward precision, which suits campaigns where contacting the wrong person costs money. ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["auc_pr"]}," is the right choice when positives are rare, since it focuses on the positive class in a way ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["auc_roc"]}," does not."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"threshold-selection","__idx":24},"children":["Threshold Selection"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["The decision threshold is chosen during training and travels with the model, which is why ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_predict"]}," needs no threshold parameter of its own."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["It is selected on ",{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["out-of-fold"]}," cross-validation predictions rather than on the model's own training scores. This matters more than it sounds. Boosted trees can nearly memorize their training split, which makes in-sample probabilities look far more separable than they really are and produces a threshold that collapses the moment it meets new data. Fitting the threshold out-of-fold avoids that. The untouched test split is then scored separately, purely for honest reporting."]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"26%","data-label":"Strategy"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Strategy"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"74%","data-label":"Picks the cutoff that"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Picks the cutoff that"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["precision_optimal"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Maximizes precision, subject to the ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["min_metric_floor"]}," guardrail. The default, and the usual choice when contacting a customer costs money."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["f1_optimal"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Balances precision and recall evenly. A reasonable neutral default when you have no strong preference."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["recall_optimal"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Maximizes recall, subject to the guardrail. Use it when missing a positive is the expensive error — churn interception, fraud triage."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["youden_j"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Maximizes true positive rate minus false positive rate. A balanced, threshold-theoretic choice that ignores class prevalence."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fixed"]}]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Uses ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fixed_threshold"]}," verbatim, with no search. Use it when the cutoff is set by a business rule rather than by the data."]}]}]}]}]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["min_metric_floor"]}," is the guardrail that stops an optimizer from producing a technically-optimal but useless operating point — a threshold with perfect precision that flags eleven people. Its direction follows ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fbeta_beta"]},": when ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fbeta_beta"]}," is below 1 the floor is a minimum recall, and when it is 1 or above the floor is a minimum precision."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["For most marketing work you can ignore the label and the threshold entirely, and target by score. The threshold matters when you need one automatic yes/no decision per customer."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"reading-the-diagnostics","__idx":25},"children":["Reading the Diagnostics"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Training writes far more than a metrics row. The tables worth reading first, in order:"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_run_log"]}," against ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_train_run_log"]},"."]}," These carry the same metrics on the test and train splits. Read them side by side. A model that scores far better on train than test is overfitting, and the gap tells you how much. A model that scores similarly on both is generalizing, whatever the absolute numbers say."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_feature_importance"]}," and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_shap_directional"]},"."]}," SHAP importance ranks which features moved the model most. The directional table goes further and says which way each one pushed, with the share of customers it pushed up versus down. Read this before you activate anything. One feature dominating the ranking is the classic signature of label leakage."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":[{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_prediction_bins"]}," and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_roc_curve"]},"."]}," The score histogram shows how the population spreads across the 0–1 range. A healthy model separates; a poor one piles everyone into the middle. The ROC curve is the standard picture of the precision-recall trade-off across every possible threshold, useful when you want to argue for a different cutoff than the one selected."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["PSI"]},", recorded in the run log, compares the train and test score distributions. It is a drift measure: a large value means the two splits are scoring differently, which usually points to a time-ordered split or a population shift rather than a modeling problem."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Two headline metrics deserve a note. ",{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["AUC-ROC"]}," measures overall ranking quality and is the number most people ask for, but it flatters models on imbalanced data. ",{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["AUC-PR"]}," focuses on the positive class and is the more honest headline when positives are rare — which is why ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["aucpr"]}," is the default ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["eval_metric"]},". The run log also records ",{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["pseudo-R²"]}," (McFadden and Cox-Snell) and a chi-square statistic for readers who want a goodness-of-fit view alongside the ranking metrics."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Treat ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["score"]}," as a reliable ",{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["ranking"]}," — higher genuinely means more likely — rather than as a literal probability. Validate against real outcomes before reading a score of 0.30 as exactly a 30% chance."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"preventing-label-leakage","__idx":26},"children":["Preventing Label Leakage"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Leakage is the most common way a propensity model looks excellent in testing and fails in production. It happens when a feature is really a stand-in for the outcome: a ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["conversion_date"]}," that only exists for converters, a ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["cancellation_reason"]}," populated only for churners, an account-status flag updated at the moment the event occurred."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["The solution removes some of this risk for you. ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["session_id"]}," is dropped automatically, correlated numeric features above ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["corr_threshold"]}," are pruned, and the identity column named in ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id_column"]}," is excluded from the feature set."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["The rest is yours to catch:"]},{"$$mdtype":"Tag","name":"ul","attributes":{},"children":[{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Drop post-outcome columns"]}," with ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["columns_to_drop"]},". Anything whose value was written at or after the moment the outcome happened does not belong in the feature set."]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Drop raw identifiers and timestamps."]}," Add Treasure AI's implicit time column to ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["columns_to_drop"]}," if it survives into your feature table."]},{"$$mdtype":"Tag","name":"li","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Read the SHAP output."]}," A single feature accounting for most of the model's importance is a leak until proven otherwise. ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_target_correlations"]}," is the fast version of the same check — a feature correlating near-perfectly with the target is almost never a real signal."]}]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"model-storage","__idx":27},"children":["Model Storage"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Trained models are saved to managed model storage under ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_name"]},", unique per Treasure AI account, and retrieved automatically by ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_predict"]},"."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["The stored artifact carries more than the fitted estimator. It also holds the run configuration, the ordered feature names, the winning hyperparameters, the selected threshold, and the preprocessing metadata — category vocabularies and the list of correlated features that were dropped. That is why prediction reproduces training exactly, and why the prediction input does not have to match the training table column for column: the model knows what it expects and reshapes what it receives."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":["Reusing a ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_name"]}," overwrites the model saved under it. When comparing algorithms, give each its own name — ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["churn_xgb_v1"]}," and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["churn_cat_v1"]}," — so one training run does not silently replace another."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"faqs","__idx":28},"children":["FAQs"]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"getting-started","__idx":29},"children":["Getting Started"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["What data do I need?"]}," A flat table with one row per customer, a binary target column marking the outcome you want to predict, and any number of feature columns describing each customer. You need enough historical examples of both outcomes for the model to learn the pattern. Feature engineering happens upstream; the solution does no aggregation of its own."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Do my columns need specific names?"]}," No. ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["target_column"]}," and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id_column"]}," point at whatever your table already uses, and every remaining column is treated as a feature. This is different from RFM, which requires three fixed column names."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Which algorithm should I choose?"]}," Start with ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["xgboost"]},", the default. If your features are heavily categorical, train ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["catboost"]}," under a second ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["model_name"]}," and compare the ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_run_log"]}," rows. Both share the same evaluation and output schema, so the comparison is direct."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Is there a separate tuning step like NBA has?"]}," No. Hyperparameter search is embedded in ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_train"]},", so training and prediction are the only two functions."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"data-and-labels","__idx":30},"children":["Data and Labels"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["How should I encode the label?"]}," As a column with two distinct classes, integer or boolean. A table with only one class fails the run with a user input error rather than training a model that cannot learn anything."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["What happens to missing values?"]}," Numeric missing values pass straight through to the tree model, which handles them natively — there is no imputation step. Missing or unseen categorical values map to a ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["\"missing\""]}," token."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Do I need to one-hot encode my categorical columns?"]}," No. Both algorithms handle categorical columns natively, XGBoost through the pandas ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["category"]}," dtype and CatBoost through its own categorical handling. Leave them as strings."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Will the model still work if my positive class is very rare?"]}," Yes, within limits. Leave ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["class_balance_strategy"]}," on ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["balanced"]}," and set ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["optuna_metric"]}," to ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["auc_pr"]},", which focuses on the positive class. Very rare positives still produce noisier scores, so validate against real outcomes before acting on them at scale."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Why does my model look great in testing but do poorly live?"]}," Almost always label leakage — a feature that is really a stand-in for the outcome. Read ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_feature_importance"]}," and ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["classifier_target_correlations"]}," to spot a suspiciously dominant feature, then exclude it with ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["columns_to_drop"]}," and retrain. See ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"#preventing-label-leakage"},"children":["Preventing Label Leakage"]},"."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"tuning-and-operating","__idx":31},"children":["Tuning and Operating"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["How often should I retrain versus re-score?"]}," Retrain when your customer base or business shifts enough to change behavior, typically monthly. Re-score more often, daily or weekly, against the saved model so recently active customers get fresh scores. Splitting the two lets you tune each cadence independently, and prediction is far cheaper than training."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["How do I interpret the score?"]}," As a ranking of likelihood, from 0.0 to 1.0, where higher means more likely to take the action. For campaigns, sort by score and take the top slice — the top decile, or as many customers as your budget covers — rather than relying on ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["predicted_label"]},". See the note under ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"#reading-the-diagnostics"},"children":["Reading the Diagnostics"]}," on reading scores as absolute probabilities."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["How do I change the threshold?"]}," Set ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["threshold.strategy"]}," at training time. Use ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["recall_optimal"]}," when missing a positive is the expensive error, ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["precision_optimal"]}," when contacting the wrong person is, or ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fixed"]}," with ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fixed_threshold"]}," when a business rule sets the cutoff. The threshold is chosen during training and travels with the model, so it cannot be changed at prediction time without retraining."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Can I see why the model made a prediction?"]}," Yes. Every training run writes SHAP feature importance and a directional breakdown showing which way each feature pushed scores. Use ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["shap_max_samples"]}," to control how much data that computation uses."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Can I speed up training?"]}," Set ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["tuning.n_trials: 0"]}," to skip the hyperparameter search and train on built-in defaults. Lowering ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["max_tuning_rows"]}," or ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["cv_folds"]}," are gentler versions of the same trade."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":3,"id":"scale-and-limits","__idx":32},"children":["Scale and Limits"]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["How large a table can I train on?"]}," Loading is capped by ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["max_load_rows"]},", which defaults to 10 million rows and can be raised."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["How do I score millions of customers?"]}," Prediction distributes across workers automatically, sharded by a stable hash of the identity column. Set ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["user_id_column"]}," — without it, prediction is limited to a single worker."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Can this predict more than two outcomes?"]}," No. This is a binary classifier: one yes/no outcome, one score per customer. Multi-class problems are not yet supported. Model each outcome separately if you need several classes."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Can I get scores in real time?"]}," No. Scoring runs as scheduled batch jobs. Frequent workflow runs can get you to a sub-daily cadence, but true real-time serving is not supported."]},{"$$mdtype":"Tag","name":"p","attributes":{},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["How do I activate scores in a campaign?"]}," The output table carries a score and a ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["score_bucket"]}," for every scored customer. Join it to your customer table so it becomes a Master Segment attribute, then build segments from it — a high-intent segment as ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["score >= 0.7"]},", or a top-decile audience — and activate through standard CDP activation workflows."]},{"$$mdtype":"Tag","name":"Heading","attributes":{"level":2,"id":"glossary","__idx":33},"children":["Glossary"]},{"$$mdtype":"Tag","name":"div","attributes":{"className":"md-table-wrapper"},"children":[{"$$mdtype":"Tag","name":"table","attributes":{"className":"md"},"children":[{"$$mdtype":"Tag","name":"thead","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"th","attributes":{"width":"24%","data-label":"Term"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Term"]}," "]},{"$$mdtype":"Tag","name":"th","attributes":{"width":"76%","data-label":"Definition"},"children":[{"$$mdtype":"Tag","name":"strong","attributes":{},"children":["Definition"]}," "]}]}]},{"$$mdtype":"Tag","name":"tbody","attributes":{},"children":[{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Binary classification"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Predicting one of two outcomes for each record — will convert versus will not."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Propensity score"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The model's estimate of how likely a customer is to take the target action, from 0.0 to 1.0. Written to the ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["score"]}," column."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Target (label)"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The known outcome the model learns from, named by ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["target_column"]},". Must have at least two distinct classes."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Feature"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["An input column describing a customer — tenure, purchase count, region, RFM quartile — that the model uses to predict the target."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Class imbalance"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["When one outcome is far rarer than the other, for example 2% churn. Handled by ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["class_balance_strategy"]},"."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Threshold"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The score cutoff that turns a score into a 0/1 label. Chosen during training and stored with the model. See ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"#threshold-selection"},"children":["Threshold Selection"]},"."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Out-of-fold prediction"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["A prediction made for a row by a model that did not train on it. Used here to select the threshold honestly."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["AUC-ROC"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Area under the ROC curve. Overall ranking quality; 0.5 is random, 1.0 is perfect."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["AUC-PR"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Area under the precision-recall curve. A better headline than AUC-ROC when the positive class is rare."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Precision"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Of the customers the model flagged positive, the share that truly were."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Recall"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Of all truly positive customers, the share the model flagged."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["F-beta"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["A weighted blend of precision and recall. ",{"$$mdtype":"Tag","name":"code","attributes":{},"children":["fbeta_beta"]}," below 1 favors precision, above 1 favors recall."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["PSI"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Population Stability Index. Compares two score distributions; used here to detect drift between the train and test splits."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Pseudo-R²"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["McFadden and Cox-Snell goodness-of-fit measures for classification, reported alongside the ranking metrics."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["SHAP"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["SHapley Additive exPlanations. Attributes each prediction to its features, showing which factors pushed a customer's score up or down."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["VIF"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Variance Inflation Factor. Measures how much a numeric feature is explained by the others; high values indicate redundancy."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Label leakage"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["When a feature encodes the outcome itself, inflating test metrics and failing in production. See ",{"$$mdtype":"Tag","name":"MarkdownLink","attributes":{"href":"#preventing-label-leakage"},"children":["Preventing Label Leakage"]},"."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Optuna"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["The open-source hyperparameter optimization library that powers the embedded tuning search."]}]},{"$$mdtype":"Tag","name":"tr","attributes":{},"children":[{"$$mdtype":"Tag","name":"td","attributes":{},"children":["XGBoost / CatBoost"]},{"$$mdtype":"Tag","name":"td","attributes":{},"children":["Gradient-boosted decision-tree libraries. Both deliver high accuracy on tabular customer data; CatBoost specializes in categorical features."]}]}]}]}]}]},"headings":[{"value":"Propensity Scoring (Binary Classifier)","id":"propensity-scoring-binary-classifier","depth":1},{"value":"Overview and Use Cases","id":"overview-and-use-cases","depth":2},{"value":"How It Fits with the Other AI Signals","id":"how-it-fits-with-the-other-ai-signals","depth":3},{"value":"Quick Start","id":"quick-start","depth":2},{"value":"Train","id":"train","depth":3},{"value":"Predict","id":"predict","depth":3},{"value":"Model Configuration","id":"model-configuration","depth":2},{"value":"classifier_train","id":"classifier_train","depth":3},{"value":"Data Preparation","id":"data-preparation","depth":4},{"value":"Workflow Parameters","id":"workflow-parameters","depth":4},{"value":"Output","id":"output","depth":4},{"value":"classifier_predict","id":"classifier_predict","depth":3},{"value":"Data Preparation","id":"data-preparation-1","depth":4},{"value":"Workflow Parameters","id":"workflow-parameters-1","depth":4},{"value":"Output","id":"output-1","depth":4},{"value":"Model Types","id":"model-types","depth":2},{"value":"XGBoost ( xgboost )","id":"xgboost--xgboost-","depth":3},{"value":"Model-Specific Parameters","id":"model-specific-parameters","depth":4},{"value":"Workflow Example","id":"workflow-example","depth":4},{"value":"CatBoost ( catboost )","id":"catboost--catboost-","depth":3},{"value":"Model-Specific Parameters","id":"model-specific-parameters-1","depth":4},{"value":"Workflow Example","id":"workflow-example-1","depth":4},{"value":"Relevant Topics","id":"relevant-topics","depth":2},{"value":"Hyperparameter Tuning","id":"hyperparameter-tuning","depth":3},{"value":"Threshold Selection","id":"threshold-selection","depth":3},{"value":"Reading the Diagnostics","id":"reading-the-diagnostics","depth":3},{"value":"Preventing Label Leakage","id":"preventing-label-leakage","depth":3},{"value":"Model Storage","id":"model-storage","depth":3},{"value":"FAQs","id":"faqs","depth":2},{"value":"Getting Started","id":"getting-started","depth":3},{"value":"Data and Labels","id":"data-and-labels","depth":3},{"value":"Tuning and Operating","id":"tuning-and-operating","depth":3},{"value":"Scale and Limits","id":"scale-and-limits","depth":3},{"value":"Glossary","id":"glossary","depth":2}],"frontmatter":{"seo":{"title":"Propensity Scoring (Binary Classifier)","description":"Score every customer with the probability they will convert, churn, or respond, using gradient-boosted models on AI Signals, Treasure AI's ML platform."}},"lastModified":"2026-09-10T05:13:56.000Z","pagePropGetterError":{"message":"","name":""}},"slug":"/ja/products/customer-data-platform/machine-learning/ai-signals/propensity-ai-signals","userData":{"isAuthenticated":false,"teams":["anonymous"]},"isPublic":true}