机器学习算法（二）: 基于XGBoost的分类预测

2023年6月30日下午3:43 • 人工智能 • 阅读 79

阿里云机器学习案例（二）

1.实验室介绍

1.1 XGBoost介绍

XGBoost是2016年由华盛顿大学陈天奇老师带领开发的一个可扩展机器学习系统。严格意义上讲XGBoost并不是一种模型，而是一个可供用户轻松解决分类、回归或排序问题的软件包。它内部实现了梯度提升树(GBDT)模型，并对模型中的算法进行了诸多优化，在取得高精度的同时又保持了极快的速度，在一段时间内成为了国内外数据挖掘、机器学习领域中的大规模杀伤性武器。

更重要的是，XGBoost在系统优化和机器学习原理方面都进行了深入的考虑。毫不夸张的讲，XGBoost提供的可扩展性，可移植性与准确性推动了机器学习计算限制的上限，该系统在单台机器上运行速度比当时流行解决方案快十倍以上，甚至在分布式系统中可以处理十亿级的数据。

XGBoost的主要优点：

1.简单易用。相对其他机器学习库，用户可以轻松使用XGBoost并获得相当不错的效果。
2.高效可扩展。在处理大规模数据集时速度快效果好，对内存等硬件资源要求不高。
3.鲁棒性强。相对于深度学习模型不需要精细调参便能取得接近的效果。
4.XGBoost内部实现提升树模型，可以自动处理缺失值。

XGBoost的主要缺点：

1.相对于深度学习模型无法对时空位置建模，不能很好地捕获图像、语音、文本等高维数据。
2.在拥有海量训练数据，并能找到合适的深度学习模型时，深度学习的精度可以遥遥领先XGBoost。

1.2 XGBoost的应用

XGBoost在机器学习与数据挖掘领域有着极为广泛的应用。据统计在2015年Kaggle平台上29个获奖方案中，17只队伍使用了XGBoost；在2015年KDD-Cup中，前十名的队伍均使用了XGBoost，且集成其他模型比不上调节XGBoost的参数所带来的提升。这些实实在在的例子都表明，XGBoost在各种问题上都可以取得非常好的效果。

同时，XGBoost还被成功应用在工业界与学术界的各种问题中。例如商店销售额预测、高能物理事件分类、web文本分类;用户行为预测、运动检测、广告点击率预测、恶意软件分类、灾害风险预测、在线课程退学率预测。虽然领域相关的数据分析和特性工程在这些解决方案中也发挥了重要作用，但学习者与实践者对XGBoost的一致选择表明了这一软件包的影响力与重要性。

2.实验室手册

2.1 学习目标

2.1.1 了解 XGBoost 的参数与相关知识

2.1.2 掌握 XGBoost 的Python调用并将其运用到天气数据集预测

2.2 代码流程

Part1 基于天气数据集的XGBoost分类实践

Step1: 库函数导入
Step2: 数据读取/载入
Step3: 数据信息简单查看
Step4: 可视化描述
Step5: 对离散变量进行编码
Step6: 利用 XGBoost 进行训练与预测
Step7: 利用 XGBoost 进行特征选择
Step8: 通过调整参数获得更好的效果

2.3 算法实战

2.3.1 基于天气数据集的XGBoost分类实战

%%%数据集地址如下%%% ：

https://tianchi-media.oss-cn-beijing.aliyuncs.com/DSW/7XGBoost/train.csv

step1:库函数导入

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
import seaborn as sns
import warnings
import xgboost

step2:数据载入

plt.rcParams['axes.unicode_minus'] = False

warnings.filterwarnings('ignore')
data = pd.read_csv(r"E:\train.csv")

step3:数据信息简单查看

data.info
data.head()
data = data.fillna(-1)
data.tail()
pd.Series(data['RainTomorrow']).value_counts()
data.describe()

step4:可视化描述

numerical_features = [x for x in data.columns if data[x].dtype == np.float64]
category_features = [x for x in data.columns if data[x].dtype != np.float64 and x != 'RainTomorrow']

sns.pairplot(data = data[['Rainfall',
                          'Evaporation',
                          'Sunshine'] + ['RainTomorrow']], diag_kind = 'hist', hue = 'RainTomorrow')
plt.show()
for col in data[numerical_features].columns:
    if col != 'RainTomorrow':
        sns.boxplot(x = 'RainTomorrow', y = col, saturation = 0.5, palette = 'pastel', data = data)
        plt.title(col)
        plt.show()

tlog = {}
for i in category_features:
    tlog[i] = data[data['RainTomorrow'] == 'Yes'][i].value_counts()
flog = {}
for i in category_features:
    flog[i] = data[data['RainTomorrow'] == 'No'][i].value_counts()
plt.figure(figsize=(10, 10))
plt.subplot(1,2,1)
plt.title('RainTomorrow')
sns.barplot(x = pd.DataFrame(tlog['Location']).sort_index()['Location'], y = pd.DataFrame(tlog['Location']).sort_index().index, color = 'red')
plt.subplot(1,2,2)
plt.title('Not RainTomorrow')
sns.barplot(x = pd.DataFrame(flog['Location']).sort_index()['Location'], y = pd.DataFrame(flog['Location']).sort_index().index, color = 'blue')
plt.show()
plt.figure(figsize=(10,2))
plt.subplot(1,2,1)
plt.title('RainTomorrow')
sns.barplot(x = pd.DataFrame(tlog['RainToday'][:2]).sort_index()['RainToday'], y = pd.DataFrame(tlog['RainToday'][:2]).sort_index().index, color = 'red')
plt.subplot(1,2,2)
plt.title('Not RainTomorrow')
sns.barplot(x = pd.DataFrame(flog['RainToday'][:2]).sort_index()['RainToday'], y = pd.DataFrame(flog['RainToday'][:2]).sort_index().index, color = 'blue')
plt.show()

step5:对离散变量进行编码


def get_mapfunction(x):
    mapp = dict(zip(x.unique().tolist(),
                  range(len(x.unique().tolist()))))
    def mapfunction(y):
        if y in mapp:
            return mapp[y]
        else:
            return -1
    return mapfunction
for i in category_features:
    data[i] = data[i].apply(get_mapfunction(data[i]))
data['RainTomorrow'] = data['RainTomorrow'].apply(get_mapfunction(data['RainTomorrow']))

step6:接下来利用XGBoost进行训练与预测


from sklearn.model_selection import train_test_split

data_features_part = data[[x for x in data.columns if x != 'RainTomorrow']]
data_target_part = data['RainTomorrow']

x_train, x_test, y_train, y_test = train_test_split(data_features_part, data_target_part, test_size=0.2, random_state= 2020)

from xgboost.sklearn import XGBClassifier

clf = XGBClassifier()

clf.fit(x_train, y_train)

train_predict = clf.predict(x_train)
test_predict = clf.predict(x_test)
from sklearn import metrics

print('The accuracy of the Logistic Regression is:', metrics.accuracy_score(y_train, train_predict))
print('The accuracy of the Logistic Regression is:', metrics.accuracy_score(y_test, test_predict))

confusion_matrix_result = metrics.confusion_matrix(test_predict, y_test)
print('The confusion matrix result:\n', confusion_matrix_result)

plt.figure(figsize=(8,6))
sns.heatmap(confusion_matrix_result, annot=True, cmap='Blues')
plt.xlabel('Predicted labels')
plt.ylabel('True labels')
plt.show()

step7:利用XGBoost进行特征选择

?sns.barplot

sns.barplot(y=data_features_part.columns, x=clf.feature_importances_)

from sklearn.metrics import accuracy_score
from xgboost import plot_importance

def estimate(model, data):

    ax1 = plot_importance(model, importance_type='gain')
    ax1.set_title('gain')
    ax2 = plot_importance(model, importance_type='weight')
    ax2.set_title('weight')
    ax3 = plot_importance(model, importance_type='cover')
    ax3.set_title('cover')
    plt.show()

def classes(data, label, test):
    model = XGBClassifier()
    model.fit(data, label)
    ans = model.predict(test)
    estimate(model, data)
    return ans

ans = classes(x_train, y_train, x_test)
pre = accuracy_score(y_test, ans)
print('acc=', pre)

step8:通过调整参数获得更好的效果


from sklearn.model_selection import GridSearchCV

learning_rate = [0.1, 0.3, 0.6]
subsample = [0.8, 0.9]
colsample_bytree = [0.6, 0.8]
max_depth = [3, 6, 8]

parameters = {
    'learning_rate': learning_rate,
    'subsample': subsample,
    'colsample_bytree': colsample_bytree,
    'max_depth': max_depth
}
model = XGBClassifier(n_estimators = 50)

clf = GridSearchCV(model, parameters, cv=3, scoring='accuracy', verbose=1, n_jobs=-1)
clf = clf.fit(x_train, y_train)
clf.best_params_

clf = XGBClassifier(colsample_bytree = 0.8, learning_rate = 0.3, max_depth = 6, subsample = 0.8)

clf.fit(x_train, y_train)

train_predict = clf.predict(x_train)
test_predict = clf.predict(x_test)

print('The accuracy of the Logistic Regression is:', metrics.accuracy_score(y_train, train_predict))
print('The accuracy of the Logistic Regression is:', metrics.accuracy_score(y_test, test_predict))

confusion_matrix_result = metrics.confusion_matrix(test_predict, y_test)
print('The confusion matrix result:\n', confusion_matrix_result)

plt.figure(figsize=(8,6))
sns.heatmap(confusion_matrix_result, annot=True, cmap='Blues')
plt.xlabel('Predicted labels')
plt.ylabel('True labels')
plt.show()

3.知识点

3.1 XGBoost的重要参数

3.2 XGBoost原理粗略讲解

Original: https://blog.csdn.net/lele_god/article/details/124481995
Author: lele_god
Title: 机器学习算法（二）: 基于XGBoost的分类预测

原创文章受到原创版权保护。转载请注明出处：https://www.johngo689.com/661568/

转载文章受原作者版权保护。转载请注明原作者出处！

人工智能

【自取】最近整理的，有需要可以领取学习：

Linux核心资料大放送~

全栈面试题汇总（持续更新&可下载）

一个提高学习100%效率的工具！

【超详细】深度学习面试题目！

LeetCode Python刷题答案下载！

LeetCode Java版刷题答案下载！

LeetCode C++ 版本，抓紧保存！

LeetCode GO语言刷题答案下载！

语义分割分布式训练小结

借鉴文档https://blog.csdn.net/weixin_44966641/article/details/121872773https://zhuanlan.zhihu….

人工智能 2023年7月23日
0049
【云原生】什么是云原生？如何学习云原生？一篇文章带你了解云原生

云原生，相信这个名词大家并不陌生；云原生在近期可谓是爆火，伴随云计算的滚滚浪潮，云原生(CloudNative)的概念应运而生，云原生很火，火得一塌糊涂。可是现在很多人还是不知道什…

人工智能 2023年5月30日
0072
华为海思新品SD3403

啊哦~你想找的内容离你而去了哦内容不存在，可能为如下原因导致： ① 内容还在审核中 ② 内容以前存在，但是由于不符合新的规定而被删除 ③ 内容地址错误 ④ 作者删除了内容。可…

人工智能 2023年6月25日
0051
Windows 下安装CUDA和CUDNN以及验证是否安装成功

一、CUDA和CUDNN安装参见下面一篇博客：深度学习之CUDA+CUDNN详细安装教程二、验证是否安装成功首先验证CUDA，win+R进入CMD，在命令行输入nvcc -V…

人工智能 2023年5月23日
0092
【知识图谱】实践篇——基于知识图谱的《红楼梦》人物关系可视化及问答系统实践：part2知识获取与图谱构建、服务搭建

前序文章：【知识图谱】实践篇——基于知识图谱的《红楼梦》人物关系可视化及问答系统实践：part1项目介绍与环境准备 ; 知识获取与图谱构建其中原项目提供了关系数据如下：其中五列…

人工智能 2023年6月10日
00124
51单片机DAC数模转换

51单片机DAC数模转换 DAC介绍 1、DAC简介 DAC（Digital to analog converter）即数字模拟转换器，它可以将数字信号转换为模拟信号。它的功能与…

人工智能 2023年6月27日
0076
opencv圆形网格提取函数findCirclesGrid源码笔记

opencv–findCircle源码笔记函数处理流程源码分析 * findCirclesGrid源码 findCirclesGrid2 函数源码 – …

人工智能 2023年5月26日
0063
cv2的函数没有代码提示，最详细解决办法（重启项目版），修改-ini_.py没效果

pycharm cv2 Cannot find declaration to go to numpy可以查看源代码，但cv2不可以查看源代码 cv2没有代码提示 ctrl+左键无法…

人工智能 2023年6月18日
0094
实时通信服务中的语音解混响算法实践

导读：随着音视频交流会议的日益普及，与会者在不同的环境中遇到了越来越明显和不同的混响场景，如会议室场景、玻璃会议室场景和隔音材料较差的小房间。为了确保更好的收听清晰度和舒适性，通…

人工智能 2023年5月25日
0069
ENVI图像处理（6）：NDVI和植被指数

NDVI NDVI 植被指数 ENVI操作 * NDVI band math quick stat统计图 NDVI 定义：NDVI（Normalized Difference Ve…

人工智能 2023年6月17日
0080
【深度学习】——利用pytorch搭建一个完整的深度学习项目（构建模型、加载数据集、参数配置、训练、模型保存、预测）

目录一、深度学习项目的基本构成二、实战（猫狗分类） 1、数据集下载 2、dataset.py文件 3、model.py 4、config.py 5、predict.py 一、深…

人工智能 2023年7月22日
0075
DJL-Java开发者动手学深度学习之使用Softmax进行分类

分类在之前的文章中，我们介绍了线性回归的基本概念 DJL-Java开发者动手学深度学习之线性回归，并使用高层API实两简单线性回归模型 DJL-Java动手学深度学习之线性回归实…

人工智能 2023年7月2日
0076
tf的学习笔记

本学习笔记基于北京大学-tensorflow2.0教学视频所写，方便自己日后复习。虚拟环境创建虚拟环境：conda create -n 名称 python=3.7 进入虚拟环境…

人工智能 2023年5月25日
0053
硬件里的玄乎事

系列文章目录 1.元件基础2.电路设计3.PCB设计4.元件焊接5.板子调试6.程序设计7.算法学习8.编写exe9.检测标准10.项目举例11.职业规划文章目录前言 1、一碰…

人工智能 2023年6月29日
0090
python实现 logistic 回归二分类算法（通俗讲解逻辑回归本质与由来）

logistic回归将数据样本看作是欧式空间的点，尝试找到一个超平面，将空间分成两部分，如果样本点在”正面”，则它被分为0类；如果样本点在”负…

人工智能 2023年6月16日
0064
SVM的核函数详解

文章目录 1、核函数背景 * 核函数正式定义 2、高斯核函数 * 2.2 参数带宽σ \sigma σ的影响 2.3高斯核函数的实际意义 2、多项式核函数 4、参考资料 1、核函数…

人工智能 2023年6月15日
0093

2024 年 4 月
一	二	三	四	五	六	日
1	2	3	4	5	6	7
8	9	10	11	12	13	14
15	16	17	18	19	20	21
22	23	24	25	26	27	28
29	30