Differences

This shows you the differences between two versions of the page.

--- data_mining:xgboost [2019/05/02 22:52] – phreazer
+++ data_mining:xgboost [2020/08/02 14:12] (current) – phreazer
@@ Line 1: / Line 1: @@
 ====== XGBoost ======
 //Extreme Gradient Boosting//
 Literature: Greedy Function Approximation: A Gradient Boosting Machine, by Friedman
@@ Line 19: / Line 20: @@
 $$
-$F$ is space of functions containing all regression trees
+===== Gradient boosting =====
-$K$ is number of trees
-$f_k(x_i)$ is regression tree that maps a attribute to a score
+  * $F$ is space of functions containing all regression trees
+  * $K$ is number of trees
+  * $f_k(x_i)$ is regression tree that maps a attribute to a score
 Learn functions (trees) instead of weights in $R^d$.
@@ Line 32: / Line 35: @@
 Learning objective:
-  * Training loss: Fit of the functions to the points
+  * **Training loss**: Fit of the functions to the points
-  * Regularization: Complexity of function; Number of splitting points, l2 norm of height in each segment
+  * **Regularization**: Complexity of function; Number of splitting points, l2 norm of height in each segment
 Objective:
@@ Line 56: / Line 59: @@
   * Logistic loss $l(y_i,\hat{y}_i)=y_i \ln(1+e^{-\hat{y}_i})+(1-y_i)\ln(1+e^{\hat{y}_i})$ (LogitBoost)
-Stochastic Gradient Descent can not be applied, since trees are used.
+Stochastic Gradient descent can not be applied, since trees are used.
 Solution is **additive training**: Start with constant prediction, add a new function each time.
@@ Line 79: / Line 82: @@
-Taylor expansion approximation of loss
+==== Taylor expansion ====
 Use taylor expansion to approximate a function through a power series (polynom).
@@ Line 92: / Line 95: @@
 $$\sum^n_{i=1} [l(y_i,\hat{y}_i^{(t-1)}) + g_if_t(x_i) + \frac{1}{2}h_if_t^2(x_i)]$$ with $g_i=\delta_{\hat{y}^{(t-1)}} l(y_i,\hat{y}^{(t-1)})$ and $h_i=\delta^2_{\hat{y}^{(t-1)}} l(y_i,\hat{y}^{(t-1)})$
-With removed constants
+With removed constants (and square loss)
 $$\sum^n_{i=1} [g_if_t(x_i) + \frac{1}{2}h_if_t^2(x_i)] + \Omega(f_t)$$
 So that learning function only influences $g_i$ and $h_i$ while rest stays the same.