Python数据分析–Numpy常用函数介绍(4)–Numpy中的线性关系和数据修剪压缩

2023年5月24日上午12:46 • Python • 阅读 105

摘要：总结股票均线计算原理–线性关系，也是以后大数据处理的基础之一，NumPy的 linalg 包是专门用于线性代数计算的。作一个假设，就是一个价格可以根据N个之前的价格利用线性模型计算得出。

在上一篇文章中，在计算移动平均和指数平均时，计算了不同的权重，例如

[En]

In the previous article, when calculating moving averages and exponential averages, different weights were calculated, such as

Python数据分析--Numpy常用函数介绍(4)--Numpy中的线性关系和数据修剪压缩

和

相关权重是根据不同的计算方法计算出来的，股价可以用之前股价的线性组合来表示，即股价等于之前股价乘以各自系数的结果，但这些系数需要我们确定，也就是线性相关的权重。

[En]

The relevant weights are calculated according to different calculation methods, and a stock price can be expressed by a linear combination of the previous stock price, that is, the stock price is equal to the result of multiplying the previous stock price by their respective coefficients, but, these coefficients need us to determine, that is, a linearly related weight.

一、用线性模型预测价格
创建步骤如下：
1）先获取一个包含N个收盘价的向量（数组）：

N=10
#N=len(close)
new_close = close[-N:]
new_closes= new_close[::-1]
print (new_closes)

运行结果：[39.96 38.03 38.5  38.6  36.89 37.15 36.61 37.21 36.98 36.47]2)初始化一个N×N的二维数组 A ，元素全部为 0

A = np.zeros((N, N), float)
print ("Zeros N by N", A)

undefined

3）用数组new_closes的股价填充数组A

for i in range(N):
    A[i,] = close[-N-i-1: -1-i]
print( "A", A)

试一下运行结果，并观察填充后的数组A

4）选取合适的权重

Weights [0.11405072 0.14644403 0.18803785 0.24144538 0.31002201]和The weights : [0.2 0.2 0.2 0.2 0.2]哪一种权重更合理？用线性代数的术语来说，就是解一个最小二乘法的问题。

要确定线性模型中的权重系数，就是解决最小平方和的问题，可以使用 linalg包中的 lstsq 函数来完成这个任务

(x, residuals, rank, s) = np.linalg.lstsq(A,new_closes)

其中，x是由A,new_closes通过np.linalg.lstsq（）函数，即生成的权重（向量），residuals为残差数组、rank为A的秩、s为A的奇异值。

5)预测股价，用NumPy中的 dot()函数计算系数向量与最近N个价格构成的向量的点积（dot product）,这个点积就是向量new_closes中价格的线性组合，系数由向量 x 提供

print( np.dot(new_closes, x))

完整代码如下：

import numpy as np
from datetime import datetime
import matplotlib.pyplot as plt

def datestr2num(s): #定义一个函数
    return datetime.strptime(s.decode('ascii'),"%Y-%m-%d").date().weekday()

dates, opens, high, low, close,vol=np.loadtxt('data.csv',delimiter=',', usecols=(1,2,3,4,5,6),
                       converters={1:datestr2num},unpack=True)

N=10
#N=len(close)
new_close = close[-N:]
new_closes= new_close[::-1]

A = np.zeros((N, N), float)

for i in range(N):
    A[i,] = close[-N-i-1: -1-i]

print( "A", A)

(x, residuals, rank, s) = np.linalg.lstsq(A,new_closes)
print(x) #权重系数向量

print('\n')
print(residuals)  #残差数组
print('\n')
print(rank) #A的秩
print(s)
print('\n')#奇异值
print( np.dot(new_closes, x))

运行结果如下：

二、趋势线

趋势线是根据股价图表上许多所谓的轴心点绘制的曲线。描述价格变化的趋势。你可以让电脑以一种非常简单的方式画出趋势线。

[En]

The trend line is a curve drawn according to many so-called pivot points on the stock price chart. Describe the trend of price changes. You can let the computer draw the trend line in a very easy way.

(1) 确定枢轴点的位置。假定枢轴点位置为最高价、最低价和收盘价的算术平均值。pivots = (high + low + close ) / 3

从支点出发，可以推算出股价的所谓阻力位和支撑位。阻力位是指股价上涨时遇到的阻力，下跌前的最高价；支撑位是指股价下跌时的最低价，反弹前的最低价(阻力位和支撑位不是客观的，它们只是一个估计值)。基于这些估计值，可以画出阻力位和支撑位的趋势线。我们将当天的价格范围定义为最高价格和最低价格之间的差额。

[En]

Starting from the pivot point, the so-called resistance level and support level of stock price can be deduced. The resistance level refers to the resistance encountered when the stock price rises and the highest price before the fall; the support level refers to the lowest price when the stock price falls and the lowest price before the rebound (resistance level and support level are not objective, they are just an estimator). Based on these estimators, the trend lines of resistance level and support level can be drawn. We define the price range of the day as the difference between the highest price and the lowest price.

(2) 定义一个函数用直线 y= at + b 来拟合数据，该函数应返回系数 a 和 b，再次用到 linalg 包中的 lstsq 函数。将直线方程重写为 y = Ax 的形式，其中 A = [t 1] ， x = [a b] 。使用 ones_like 和 vstack 函数来构造数组 A

numpy.ones_like(a, dtype=None, order=’K’, subok=True) 返回与指定数组具有相同形状和数据类型的数组，并且数组中的值都为1。

numpy.vstack(tup) [source] 垂直(行)按顺序堆叠数组。这等效于形状(N,)的1-D数组已重塑为(1,N)后沿第一轴进行concatenation。重建除以vsplit的数组。如下两小例：

>>> a = np.array([1, 2, 3]) >>> b = np.array([2, 3, 4]) >>> np.vstack((a,b)) array([[1, 2, 3],               [2, 3, 4]])

>>> a = np.array([[1], [2], [3]])
>>> b = np.array([[2], [3], [4]])
>>> np.vstack((a,b))
array([[1],
       [2],
       [3],
       [2],
       [3],
       [4]])

完整代码如下：

import numpy as np
from datetime import datetime
import matplotlib.pyplot as plt

def datestr2num(s): #定义一个函数
    return datetime.strptime(s.decode('ascii'),"%Y-%m-%d").date().weekday()

dates, opens, high, low, close,vol=np.loadtxt('data.csv',delimiter=',', usecols=(1,2,3,4,5,6),
                       converters={1:datestr2num},unpack=True)
"""
N=10
#N=len(close)
new_close = close[-N:]
new_closes= new_close[::-1]

A = np.zeros((N, N), float)

for i in range(N):
    A[i,] = close[-N-i-1: -1-i]

print( "A", A)
(x, residuals, rank, s) = np.linalg.lstsq(A,new_closes)
print(x) #权重系数向量
print(residuals)  #残差数组
print(rank) #A的秩
print(s)
print( np.dot(new_closes, x))
"""
pivots = (high + low + close ) / 3

def fit_line(t, y):
    A = np.vstack([t, np.ones_like(t)]).T
np.ones_like(t) 即定义一个像t一样，有相同形状和数据类型的数组，并且数组中的值都为1
    return np.linalg.lstsq(A, y)[0]

t = np.arange(len( close)) #按close数列创建一个数列t

sa, sb = fit_line(t, pivots - (high - low)) #用直线y=at+b来拟合数据，该函数应返回系数a(sa) 和 b(sb)
ra, rb = fit_line(t, pivots + (high - low))
support = sa * t + sb     #计算支撑线数列
resistance = ra * t + rb  #计算阻力线数列

condition = (close > support) & (close < resistance)#设置一个判断数据点是否位于趋势线之间的条件，作为 where 函数的参数
between_bands = np.where(condition)

plt.plot(t, close,color='r')
plt.plot(t, support,color='g')
plt.plot(t, resistance,color='y')
plt.show()

运行结果：

三、数组的修剪和压缩

NumPy中的 ndarray 类定义了许多方法，可以对象上直接调用。通常情况下，这些方法会返回一个数组。

ndarray 对象的方法相当多，像前面遇到的 var 、 sum 、 std 、 argmax 、argmin 以及 mean 函数也均为 ndarray 方法。下面介绍一下数组的修前与压缩。

1、 clip 方法返回一个修剪过的数组：将所有比给定最大值还大的元素全部设为给定的最大值，而所有比给定最小值还小的元素全部设为给定的最小值

a = np.arange(10)
print("a =", a)
print("Clipped", a.clip(3, 7))

运行结果：

a = [0 1 2 3 4 5 6 7 8 9]
Clipped [3 3 3 3 4 5 6 7 7 7]

很明显，a.clip(3,7)将数组a中的小于3的设置为3，大于7的全部设置为7.

2、 compress 方法返回一个根据给定条件筛选后的数组

b = np.arange(10)
print (a)
print ("Compressed", a.compress(a >3))

运行结果：

[0 1 2 3 4 5 6 7 8 9]
Compressed [4 5 6 7 8 9]

四、阶乘

prod() 方法，可以计算数组中所有元素的乘积.

c = np.arange(1,5)
print("b =", c)
print("Factorial", c.prod())

运行结果：

b = [1 2 3 4]
Factorial 24

如果想知道1~8的所有阶乘值，调用 cumprod()方法，计算数组元素的累积乘积。

print( "Factorials", c.cumprod())

运行结果：

Factorials [  1   2   6  24 120]

本篇主要介绍了一个通过现在有数据，用函数 y= at + b 来拟合数据进行线性拟合后，用 linalg包中的 lstsq 函数来完成最小二乘相关后，预测股价的实例，来了解了一些numpy的函数及作用；同时介绍了数据修剪及压缩和阶乘的计算。

Original: https://www.cnblogs.com/codingchen/p/16304038.html
Author: PursuitingPeak
Title: Python数据分析–Numpy常用函数介绍(4)–Numpy中的线性关系和数据修剪压缩

原创文章受到原创版权保护。转载请注明出处：https://www.johngo689.com/499501/

转载文章受原作者版权保护。转载请注明原作者出处！

python

【自取】最近整理的，有需要可以领取学习：

Linux核心资料大放送~

全栈面试题汇总（持续更新&可下载）

一个提高学习100%效率的工具！

【超详细】深度学习面试题目！

LeetCode Python刷题答案下载！

LeetCode Java版刷题答案下载！

LeetCode C++ 版本，抓紧保存！

LeetCode GO语言刷题答案下载！

python学习_PIL的Image模块初步使用

Pillow 是 Python 中较为基础的图像处理库，主要用于图像的基本处理，比如裁剪图像、调整图像大小和图像颜色处理等。与 Pillow 相比，OpenCV 和 Scikit-…

Python 2023年10月29日
0039
使用 Kubeadm 部署 Kubernetes(K8S) 安装

1. 安装要求在开始之前，部署Kubernetes集群机器需要满足以下几个条件：一台或多台机器，操作系统 CentOS7.x-86_x64 硬件配置：2GB或更多RAM，2个C…

Python 2023年10月19日
0078
np.stack()使用

1.官网https://numpy.org/doc/stable/reference/generated/numpy.stack.html np.stack(arrays , ax…

Python 2023年8月28日
0048
科学计算库Numpy基础&提升（理解＋重要函数讲解）

对于同样的数值计算任务，使用numpy比直接编写python代码实现优点：代码更简洁： numpy直接以数组、矩阵为粒度计算并且支持大量的数学函数，而python需要用for循…

Python 2023年11月2日
0041
python绘图条形图_用matplotlib在python中绘制漂亮的条形图

我在下面写了一个python代码来为我的数据绘制一个条形图。我调整了参数，但未能使其美观(见附图)。在 python代码如下：def plotElapsedDis(axis, jv…

Python 2023年9月6日
0047
Docker安装MySQL并使用Navicat连接

MySQL简单介绍： MySQL 是一个开放源码的关系数据库管理系统，开发者为瑞典 MySQL AB 公司。目前 MySQL 被广泛地应用在 Internet 上的大中小型网站中。…

Python 2023年10月21日
0043
X-Frame-Options

X-Frame-Options 最近在 django 的项目中做相同域名内嵌网页时，出现了 Refused to display ‘http://127.0.0.1/’ in a …

Python 2023年8月4日
0073
matplotlib ax bar color 设置ax bar的颜色、透明度、label legend

matplotlib ax bar color 设置ax bar的颜色 d = nx.degree(g1) print("网络的度分布为:{}".format(…

Python 2023年8月31日
00145
Pytest（16）随机执行测试用例pytest-random-order

前言通常我们认为每个测试用例都是相互独立的，因此需要保证测试结果不依赖于测试顺序，以不同的顺序运行测试用例，可以得到相同的结果。pytest默认运行用例的顺序是按模块和用例命名的…

Python 2023年9月12日
0039
python colorbar长度_如何改变matplotlib色标colorbar的字体大小？

所以我们只能通过读源代码colorbar.py来寻找方法。上图是colorbar.py文件的说明文档，从文档中我们可以看出，Figure.colorbar方法的实现依靠于make…

Python 2023年9月4日
00236
经典同态加密算法Paillier解读 – 原理、实现和应用

摘要随着云计算和人工智能的兴起，如何安全有效地利用数据，对持有大量数字资产的企业来说至关重要。同态加密，是解决云计算和分布式机器学习中数据安全问题的关键技术，也是隐私计算中，横跨…

Python 2023年10月9日
0057
Conda常用操作

之后遇到了新的东西会慢慢的补充以下均假设：myenv是一个名为”myenv”的虚拟环境一.最重要：寻求conda的帮助 conda -h conda l…

Python 2023年9月9日
0040
drf 视图组件

内容概要 request 对象和 response 对象 GenericAPIView 介绍基于 GenericAPIView 的 5个视图扩展类 GenericAPIView …

Python 2023年11月9日
0026
【Java 数据结构】-二叉树OJ题

作者：学Java的冬瓜博客主页：☀冬瓜的主页🌙专栏：【Java 数据结构】分享：宇宙的最可理解之处在于它是不可理解的，宇宙的最不可理解之处在于它是可理解的。——《乡村教师》主要内容…

Python 2023年9月29日
0043
基于MATLAB的图片中文字的提取及识别

基于MATLAB的图片中文字的提取及识别一．引言随着计算机科学的飞速发展，以图像为主的多媒体信息迅速成为重要的信息传递媒介，在图像中，文字信息(如新闻标题等字幕) 包含了丰富的…

Python 2023年8月1日
0072
pandas计算某列每行带有分隔符的数据中包含特定值的次数

某次做一个数据的处理，要计算用户的粉丝数量，数据集大概是这样的：传播节点微博用户id关注用户idsae26e5e3db7626dcaf6819ce5492d534″0…

Python 2023年8月18日
0054

2024 年 5 月
一	二	三	四	五	六	日
		1	2	3	4	5
6	7	8	9	10	11	12
13	14	15	16	17	18	19
20	21	22	23	24	25	26
27	28	29	30	31

Python数据分析–Numpy常用函数介绍(4)–Numpy中的线性关系和数据修剪压缩

大家都在看