【数据结构初阶】堆&&堆的实现&&堆排序&&TOP-K

2023年5月30日下午1:52 • 人工智能 • 阅读 109

大家好我是沐曦希💕

文章目录

1.前言
2.堆的概念及结构
*
2.1 堆的选择题
3.堆的实现
*
3.1 堆向下调整算法
3.2 堆向上调整算法
3.3 堆的创建
–
- 3.3.1 向下调整建堆时间复杂度
- 3.3.2 向上调整建堆时间复杂度
3.4 堆的插入
3.5 堆的删除
3.6 代码的实现
Heap.h
test.c
Heap.c
4.堆的应用
*
4.1 堆排序
4.2 TOP-K问题
5.写在最后

1.前言

普通的二叉树是不适合用数组来存储的，因为可能会存在大量的空间浪费。而完全二叉树更适合使用顺序结构存储。现实中我们通常把堆(一种二叉树)使用顺序结构的数组来存储，需要注意的是这里的堆和操作系统虚拟进程地址空间中的堆是两回事，一个是数据结构，一个是操作系统中管理内存的一块区域分段。

; 2.堆的概念及结构

如果有一个关键码的集合K = { K0，K1 ，K2 ，…，K(n-1) }把它的所有元素按完全二叉树的 顺序存储方式存储在一个一维数组中，并满足： Ki。将根节点最大的堆叫做最大堆或大根堆，根节点最小的堆叫做最小堆或小根堆。
堆的性质：

堆中某个节点的值总是不大于或不小于其父节点的值；
堆总是一棵完全二叉树。

有关堆双亲和孩子数组下标计算公式：
leftchild(左孩子) = parent 2+1(奇数)
rightchild(右孩子) = parent 2+2(偶数)
parent = (child – 1)/2

2.1 堆的选择题

1.下列关键字序列为堆的是：（）
A 100,60,70,50,32,65
B 60,70,65,50,32,100
C 65,100,70,32,50,60
D 70,65,100,32,50,60
E 32,50,100,70,65,60
F 50,100,70,65,60,32

答案是：A

3.堆的实现

3.1 堆向下调整算法

现在我们给出一个数组，逻辑上看做一颗完全二叉树。我们通过 从根节点开始的向下调整算法可以把它调整成一个小堆。 向下调整算法有一个前提：左右子树必须是一个堆，才能调整。
int array[] = {27,15,19,18,28,34,65,49,25,37};

这里进行小堆的下调


void Swap(int* p1, int* p2)
{
    int tmp = *p1;
    *p1 = *p2;
    *p2 = tmp;
}
void AdjustDown(int* a, int n, int parent)
{

    int minchild = parent * 2 + 1;

    while (minchild < n)

    {
        if (a[minchild + 1] < a[minchild] && minchild + 1 < n)

        {
            minchild++;
        }
        if (a[minchild] < a[parent])

        {
            Swap(&a[minchild], &a[parent]);
            parent = minchild;
            minchild = parent * 2 + 1;

        }
        else
            break;
    }
}

#include
void Swap(int* p1, int* p2)
{
    int tmp = *p1;
    *p1 = *p2;
    *p2 = tmp;
}
void AdjustDown(int* a, int n, int parent)
{
    int minchild = parent * 2 + 1;
    while (minchild < n)
    {
        if (a[minchild + 1] < a[minchild] && minchild + 1 < n)
        {
            minchild++;
        }
        if (a[minchild] < a[parent])
        {
            Swap(&a[minchild], &a[parent]);
            parent = minchild;
            minchild = parent * 2 + 1;
        }
        else
            break;
    }
}
int main()
{
    int arr[] = { 27,15,19,18,28,34,65,49,25,37 };
    int sz = sizeof(arr) / sizeof(arr[0]);
    AdjustDown(arr, sz, 0);
    int i = 0;
    for (i = 0; i < sz; i++)
    {
        printf("%d ", arr[i]);
    }
    printf("\n");
    return 0;
}

此时任意节点的值总是不小于其父节点的值，即为小堆。

那么要建大堆则需要通过堆向上调整算法。

3.2 堆向上调整算法

#include
void Swap(int* p1, int* p2)
{
    int tmp = *p1;
    *p1 = *p2;
    *p2 = tmp;
}
void AdjustUp(int* a, int child)
{
    int parent = (child - 1) / 2;

    while (child > 0)
    {
        if (a[child] > a[parent])
        {
            Swap(&a[child], &a[parent]);
            child = parent;
            parent = (child - 1) / 2;
        }
        else
        {
            break;
        }
    }
}
int main()
{
    int arr[] = { 27,15,19,18,28,34,65,49,25,37 };
    int sz = sizeof(arr) / sizeof(arr[0]);
    int i = 0;
    for (i = 0; i < sz; i++)
    {
        AdjustUp(arr, i);
    }
    for (i = 0; i < sz; i++)
    {
        printf("%d ", arr[i]);
    }
    printf("\n");
    return 0;
}

此时任意节点的值总是不大于其父节点的值，即为大堆。

3.3 堆的创建

下面我们给出一个数组，这个数组逻辑上可以看做一颗完全二叉树，但是还不是一个堆，现在我们通过算法，把它构建成一个堆。根节点左右子树不是堆，我们怎么调整呢？这里我们从倒数的第一个非叶子节点的子树开始调整，一直调整到根节点的树，就可以调整成堆。
int a[] = {1,5,3,8,7,6};

每插进一个结点，便调用一次堆向上调整算法或者向下调用调整算法。

; 3.3.1 向下调整建堆时间复杂度

因为堆是完全二叉树，而满二叉树也是完全二叉树，此处为了简化使用满二叉树来证明(时间复杂度本来看的就是近似值，多几个节点不影响最终结果)：

因此：向下调整建堆的时间复杂度为O(N)。

3.3.2 向上调整建堆时间复杂度

因此：向上调整建堆的时间复杂度为O(N)。

; 3.4 堆的插入

先插入一个10到数组的尾上，再进行向上调整算法，直到满足堆。

void HeapPush(HP* php, HpDataType x)
{
    assert(php);
    if (php->capacity == php->size)
    {
        int newcapacity = php->capacity == 0 ? 4 : php->capacity * 2;
        HpDataType* tmp = (HpDataType*)realloc(php->a, sizeof(HpDataType) * newcapacity);
        if (tmp == NULL)
        {
            perror("realloc fail");
            exit(-1);
        }
        php->a = tmp;
        php->capacity = newcapacity;
    }
    php->a[php->size] = x;
    php->size++;

    AdjustUp(php->a, php->size - 1);
}

3.5 堆的删除

删除堆是删除堆顶的数据，将堆顶的数据根最后一个数据一换，然后删除数组最后一个数据，再进行向下调整算法。

堆的删除的用途：得到次大或者次小的数，可以找到前K个大或者小的。

void HeapPop(HP* php)
{
    assert(php);
    assert(!HeapEmpty(php));
    Swap(&php->a[0], &php->a[php->size - 1]);
    php->size--;
    AdjustDown(php->a, php->size, 0);
}

3.6 代码的实现

Heap.h

#pragma once
#define _CRT_SECURE_NO_WARNINGS 1
#include
#include
#include
#include
typedef int HpDataType;
typedef struct Heap
{
    HpDataType* a;
    int size;
    int capacity;
}HP;

void HeapInit(HP* php);

void HeapDeStory(HP* php);

void HeapPush(HP* php, HpDataType x);

void HeapPop(HP* php);

bool HeapEmpty(HP* php);

int HeapSize(HP* php);

void HeapPrint(HP* php);

HpDataType HeapTop(HP* php);

test.c

#include"Heap.h"
int main()
{
    int arr[] = { 15,18,19,25,28,34,65,49,27,37 };
    HP hp;
    HeapInit(&hp);
    int sz = sizeof(arr) / sizeof(arr[0]);
    int i = 0;
    for (i = 0; i < sz; i++)
    {
        HeapPush(&hp, arr[i]);
    }
    printf("%d\n", HeapTop(&hp));
    HeapPrint(&hp);
    HeapPush(&hp, 10);
    printf("%d\n", HeapTop(&hp));
    HeapPrint(&hp);
    HeapPop(&hp);
    printf("%d\n", HeapTop(&hp));
    HeapPrint(&hp);
    HeapPop(&hp);
    printf("%d\n", HeapTop(&hp));
    HeapPrint(&hp);
    HeapDeStory(&hp);
    return 0;
}

Heap.c

#include"Heap.h"
void HeapInit(HP* php)
{
    assert(php);
    php->a = NULL;
    php->capacity = 0;
    php->size = 0;
}
void HeapDeStory(HP* php)
{
    assert(php);
    free(php->a);
    php->capacity = 0;
    php->size = 0;
    php->a = NULL;
}
void Swap(HpDataType* pc, HpDataType* pp)
{
    HpDataType tmp = *pc;
    *pc = *pp;
    *pp = tmp;
}
void AdjustUp(HpDataType* a, int child)
{
    int parent = (child - 1) / 2;
    while (child > 0)
    {
        if (a[child] < a[parent])
        {
            Swap(&a[child], &a[parent]);
            child = parent;
            parent = (child - 1) / 2;
        }
        else
            break;
    }
}
void HeapPush(HP* php, HpDataType x)
{
    assert(php);
    if (php->capacity == php->size)
    {
        int newcapacity = php->capacity == 0 ? 4 : php->capacity * 2;
        HpDataType* tmp = (HpDataType*)realloc(php->a, sizeof(HpDataType) * newcapacity);
        if (tmp == NULL)
        {
            perror("realloc fail");
            exit(-1);
        }
        php->a = tmp;
        php->capacity = newcapacity;
    }
    php->a[php->size] = x;
    php->size++;

    AdjustUp(php->a, php->size - 1);
}
void AdjustDown(HpDataType* a, int n, int parent)
{
    int minchild = parent * 2 + 1;
    while (minchild < n)
    {
        if (a[minchild + 1] < a[minchild] && minchild + 1 < n)
        {
            minchild++;
        }
        if (a[minchild] < a[parent])
        {
            Swap(&a[minchild], &a[parent]);
            parent = minchild;
            minchild = parent * 2 + 1;
        }
        else
            break;
    }
}
void HeapPop(HP* php)
{
    assert(php);
    assert(!HeapEmpty(php));
    Swap(&php->a[0], &php->a[php->size - 1]);
    php->size--;
    AdjustDown(php->a, php->size, 0);
}
bool HeapEmpty(HP* php)
{
    assert(php);
    return php->size == 0;
}

int HeapSize(HP* php)
{
    assert(php);
    return php->size;
}
void HeapPrint(HP* php)
{
    assert(php);
    int i = 0;
    for (i = 0; i < php->size; i++)
    {
        printf("%d ", php->a[i]);
    }
    printf("\n");
}
HpDataType HeapTop(HP* php)
{
    assert(php);
    assert(!HeapEmpty(php));
    return php->a[0];
}

4.堆的应用

4.1 堆排序

堆排序即利用堆的思想来进行排序，总共分为两个步骤：

建堆:
1.1 升序：建大堆
1.2 降序：建小堆
利用堆删除思想来进行排序:
建堆和堆删除中都用到了向下调整，因此掌握了向下调整，就可以完成堆排序。

升序建大堆，然后第一个和最后一个位置交换，把最后一个不看做堆里面的，进行向下调整处理，迭代出最大的，后续依次处理。

#include
void Swap(int* p1, int* p2)
{
    int tmp = *p1;
    *p1 = *p2;
    *p2 = tmp;
}

void AdjustDown(int* a, int n, int parent)
{
    int maxchild = parent * 2 + 1;
    while (maxchild < n)
    {
        if (a[maxchild] < a[maxchild + 1] && maxchild + 1 < n)
        {
            maxchild++;
        }
        if (a[maxchild] > a[parent])
        {
            Swap(&a[maxchild], &a[parent]);
            maxchild = parent;
            parent = (maxchild - 1) / 2;
        }
        else
            break;
    }
}
void HeapSort(int* a, int sz)
{
    int i = 0;

    for (i = (sz - 1 - 1) / 2; i >= 0; --i)
    {
        AdjustDown(a, sz, i);
    }
    for (i = 0; i < sz; i++)
    {
        printf("%d ", a[i]);
    }
    printf("\n");

    i = 1;
    while (i < sz)
    {
        Swap(&a[0], &a[sz - i]);
        AdjustDown(a, sz - 1, 0);
        i++;
    }
    for (i = 0; i < sz; i++)
    {
        printf("%d ", a[i]);
    }
    printf("\n");
}
int main()
{
    int arr[] = { 4,5,3,20,17,16 };
    int sz = sizeof(arr) / sizeof(arr[0]);
    HeapSort(arr, sz);
    return 0;
}

建堆时候，升序采用大堆，从倒数第一个非叶子节点(最后一个叶子节点的父亲)开始下调，直到调整到根。因为叶子是没有孩子和它进行上下调整，从倒数第一层开始调整。

而不采用小堆：因为根与子树排好了，剩下数据看作堆，父子关系全乱了。只能重新堆剩下数据再一次使用向下调整建堆，选出次小数据，每次时间复杂度为O(N)，效率低。

4.2 TOP-K问题

TOP-K问题：即求数据结合中前K个最大的元素或者最小的元素，一般情况下数据量都比较大。
比如：专业前10名、世界500强、富豪榜、游戏中前100的活跃玩家等。
对于Top-K问题，能想到的最简单直接的方式就是排序，但是：如果数据量非常大，排序就不太可取了(可能数据都不能一下子全部加载到内存中)。最佳的方式就是用堆来解决，基本思路如下：

用数据集合中前K个元素来建堆
1.1 前k个最大的元素，则建小堆
1.2 前k个最小的元素，则建大堆
用剩余的N-K个元素依次与堆顶元素来比较，不满足则替换堆顶元素。
将剩余N-K个元素依次与堆顶元素比完之后，堆中剩余的K个元素就是所求的前K个最小或者最大的元素。

#include
#include
#include
#include
void CreateDataFile(const char* filename, int N)
{
    assert(filename);
    FILE* fin = fopen(filename, "w");
    if (fin == NULL)
    {
        perror("fopen fail");
        exit(-1);
    }
    int i = 0;
    for (i = 0; i < N; i++)
    {
        fprintf(fin, "%d\n", rand() % 1000);
    }
    fclose(fin);
}
void Swap(int* p1, int* p2)
{
    int tmp = *p1;
    *p1 = *p2;
    *p2 = tmp;
}
void AdjustDown(int* a, int n, int parent)
{
    int minChild = parent * 2 + 1;
    while (minChild < n)
    {

        if (minChild + 1 < n && a[minChild + 1] < a[minChild])
        {
            minChild++;
        }

        if (a[minChild] < a[parent])
        {
            Swap(&a[minChild], &a[parent]);
            parent = minChild;
            minChild = parent * 2 + 1;
        }
        else
        {
            break;
        }
    }
}
void PrintTok(const char* filename, int k)
{
    assert(filename);
    FILE* fout = fopen(filename, "r");
    if (fout == NULL)
    {
        perror("fopen fail");
        exit(-1);
    }
    int i = 0;
    int* maxHeap = (int*)malloc(sizeof(int) * k);
    if (maxHeap == NULL)
    {
        perror("malloc fail");
        exit(-1);
    }

    for (i = 0; i < k; i++)
    {
        fscanf(fout, "%d", &maxHeap[i]);
    }

    for (int j = (k - 2) / 2; j >= 0; --j)
    {
        AdjustDown(maxHeap, k, j);
    }
    int val = 0;

    while (fscanf(fout, "%d", &val) != EOF)
    {
        if (val > maxHeap[0])
        {
            maxHeap[0] = val;
            AdjustDown(maxHeap, k, 0);
        }
    }
    for (i = 0; i < k; i++)
    {
        printf("%d ", maxHeap[i]);
    }
    free(maxHeap);
    maxHeap = NULL;
    fclose(fout);
}
int main()
{
    const char* filename = "Data.txt";
    srand((unsigned int)time(NULL));
    int N = 10000;
    int k = 10;

    PrintTok(filename, k);
    return 0;
}

5.写在最后

那么堆以及堆应用：堆排序和TOP-K问题就到这里了。

Original: https://blog.csdn.net/m0_68931081/article/details/126227194
Author: 沐曦希
Title: 【数据结构初阶】堆&&堆的实现&&堆排序&&TOP-K

原创文章受到原创版权保护。转载请注明出处：https://www.johngo689.com/542945/

转载文章受原作者版权保护。转载请注明原作者出处！

人工智能

【自取】最近整理的，有需要可以领取学习：

Linux核心资料大放送~

全栈面试题汇总（持续更新&可下载）

一个提高学习100%效率的工具！

【超详细】深度学习面试题目！

LeetCode Python刷题答案下载！

LeetCode Java版刷题答案下载！

LeetCode C++ 版本，抓紧保存！

LeetCode GO语言刷题答案下载！

人工智能与神经网络-它怎么工作

举个栗子：对于一张图片的处理数据输入假设我有一个图像数据，大小是64 * 64个像素一个像素就是一个颜色点，一个颜色点由红绿蓝三个值来表示，例如，红绿蓝为255,255,255…

人工智能 2023年7月13日
0064
史上最全学习率调整策略lr_scheduler

学习率是深度学习训练中至关重要的参数，很多时候一个合适的学习率才能发挥出模型的较大潜力。所以学习率调整策略同样至关重要，这篇博客介绍一下Pytorch中常见的学习率调整方法。 im…

人工智能 2023年6月12日
00109
【python数字信号处理】——DFT、DTFT（频谱图、幅度图、相位图）

目录一、离散时间傅里叶变换DTFT 二、离散傅里叶变换DFT 三、DFT与DTFT的关系参考：《数字信号处理》——（一）.DTFT、DFT(python实现)远行者223…

人工智能 2023年7月6日
0058
Python 大数据的进行信用卡欺诈检测（附源码与注释）

本案例可用于帮助大家对前面知识的掌握，同样也可以用于毕业设计等用途，我写文的初衷只是帮助大家对知识的掌握。一、背景和目的该数据集包含使用信用卡进行的金融交易的数据。这些数据是指…

人工智能 2023年7月15日
0060
轻量化网络总结[1]–SqueezeNet，Xception，MobileNetv1~v3

笔者还写了《轻量化网络总结[2]–ShuffleNetv1/v2，OSNet，GhostNet》，点击即可查看轻量化网络 * – 1. SqueezeNet &#8…

人工智能 2023年7月13日
0042
机器学习算法系列（七）-对数几率回归算法（一）（Logistic Regression Algorithm）

阅读本文需要的背景知识点：线性回归、最大似然估计、一丢丢编程知识一、引言前面几节我们学习了标准线性回归，然后介绍了三种正则化的方法 – 岭回归、Lasso回归、弹性…

人工智能 2023年6月17日
0079
关于汽车领域的知识图谱实战入门

根据https://www.bilibili.com/video/BV1iv411k7qG整理 01实体识别基于nlp的g3语言去抽取实体对象和基于关系抽取的情境下，用到命名实体…

人工智能 2023年6月1日
0064
移动边缘计算终端如何赋能高校学习空间智慧管理

人工智能时代，我们能否让图书馆的学习空间智能化，可以自主思考呢？我们的解决方案，融合了人工智能的最新软件和硬件，可以更高效地利用每个座位，迅速解决读者面临的实际问题，并赋能管理人…

人工智能 2023年6月4日
0081
使用Roberts算子进行图像分割（Matlab自编程实现）

首先了解一下何为边缘，”边缘”可定义为图像局部区域特征不相同的那些区域间的分界线，而”线”则可以认为是具有很小宽度的其中间区域具有相…

人工智能 2023年6月22日
0076
5个疯狂的 Python 项目创意

你知道 Python 是被称为全能编程语言的吗？是的，它确实是，虽然不应该在每个项目中都使用它。你可以使用它来创建桌面应用程序、游戏、移动应用程序、网站和系统软件。它甚至是最适…

人工智能 2023年5月27日
0056
【数据科学】05 数据合并（merge、concat、combine）与数据清洗（缺失值、重复值、内容和格式）

实际应用中，需要分析的数据可能来自不同的数据集，因此在开始数据分析之前，需要先将不同的数据集合并。pandas中提供了三种不同的数据合并方式： 1.1 merge()合并 pd.m…

人工智能 2023年7月6日
0087
pandas教程03—DataFrame的创建及索引

文章目录欢迎关注公众号【Python开发实战】，免费领取Python学习电子书！工具-pandas * Dataframe对象 – 创建Dataframe 多级索引…

人工智能 2023年7月6日
0081
全是狠活！SpringBoot文档也太那个了，图文并茂详尽讲解

前沿 SpringBoot是由Pivotal团队提供的在Spring框架基础之上开发的框架，其设计目的是用来简化应用的初始搭建以及开发过程。SpringBoot本身并不提供Spri…

人工智能 2023年6月27日
0070
AdaPrompt: Adaptive Prompt-based Finetuning for Relation Extraction

AdaPrompt: Adaptive Prompt-based Finetuning for Relation Extraction 本文仅供参考、交流、学习论文地址：https…

人工智能 2023年5月28日
0083
【目标检测】YOLO v5 吸烟行为识别检测

提示：文章写完后，目录可以自动生成，如何生成可参考右边的帮助文档 YOLO v5 吸烟行为目标检测模型：计算机配置、制作数据集、训练、结果分析和使用前言相关连接（look评论）…

人工智能 2023年6月23日
0089
Yolov5口罩佩戴实时检测项目（模型剪枝+opencv+python推理）

目录 0. 前言 1. 训练 * 1.1 获取口罩佩戴检测数据集 1.2 训练环境配置 1.3 修改模型文件和数据集文件 – 1.3.1 使用的模型 1.3.2 下载y…

人工智能 2023年7月19日
0073

2024 年 4 月
一	二	三	四	五	六	日
1	2	3	4	5	6	7
8	9	10	11	12	13	14
15	16	17	18	19	20	21
22	23	24	25	26	27	28
29	30