当前位置：移动技术网 > IT编程>脚本编程>Python > Python BeautifulSoup 使用

Python BeautifulSoup 使用

2019年01月21日 | 移动技术网IT编程 | 我要评论

鬼泣4automatic,潘姿彤,97bao

bs4库简单使用:

1.最好配合lxml库，下载：pip install lxml

2.最好配合requests库，下载：pip install requests

3.下载bs4：pip install bs4

4.直接输入pip没用？解决：环境变量->系统变量->path->新建：c:\python27\scripts

案例：获取网站标题

# -*- coding:utf-8 -*-
from bs4 import beautifulsoup
import requests
 
url = "https://www.baidu.com"
 
response = requests.get(url)
 
soup = beautifulsoup(response.content, 'lxml')
 
print soup.title.text

标签识别

示例1：

# -*- coding:utf-8 -*-
from bs4 import beautifulsoup
 
html = '''
<html>
<head><title>the dormouse's story</title></head>
<body>
<p class="title"><b>the dormouse's story</b></p>
<p class="story">once upon a time there were three little sisters; and their names were
<a href="http://example.com/elsie" class="sister" id="link1">elsie</a>,
<a href="http://example.com/lacie" class="sister" id="link2">lacie</a> and
<a href="http://example.com/tillie" class="sister" id="link3">tillie</a>;
and they lived at the bottom of a well.</p>
<p class="story">...</p>
</body>
</html>
'''
soup = beautifulsoup(html, 'lxml')
 
# beautifulsoup中有内置的方法来实现格式化输出
print(soup.prettify())
 
# title标签内容
print(soup.title.string)
 
# title标签的父节点名
print(soup.title.parent.name)
 
# 标签名为p的内容
print(soup.p)
 
# 标签名为p的class内容
print(soup.p["class"])
 
# 标签名为a的内容
print(soup.a)
 
# 查找所有的字符a
print(soup.find_all('a'))
 
# 查找id='link3'的内容
print(soup.find(id='link3'))

示例2：

# -*- coding:utf-8 -*-
from bs4 import beautifulsoup
 
html = '''
<html>
<head><title>the dormouse's story</title></head>
<body>
<p class="story">once upon a time there were three little sisters; and their names were
<a href="http://example.com/elsie" class="sister" id="link1">elsie</a>,
<a href="http://example.com/lacie" class="sister" id="link2">lacie</a> and
<a href="http://example.com/tillie" class="sister" id="link3">tillie</a>;
and they lived at the bottom of a well.</p>
<p class="story">...</p>
</body>
</html>
'''
 
soup = beautifulsoup(html, 'lxml')
 
# 将p标签下的所有子标签存入到了一个列表中
print (soup.p.contents)

find_all示例:

# -*- coding:utf-8 -*-
from bs4 import beautifulsoup
 
html = '''
<div class="panel">
    <div class="panel-heading">
        <h4>hello</h4>
    </div>
    <div class="panel-body">
        <ul class="list" id="list-1">
            <li class="element">foo</li>
            <li class="element">bar</li>
            <li class="element">jay</li>
        </ul>
        <ul class="list list-small" id="list-2">
            <li class="element">foo</li>
            <li class="element">bar</li>
        </ul>
    </div>
</div>
'''
 
soup = beautifulsoup(html, 'lxml')
 
# 查找所有的ul标签内容
print(soup.find_all('ul'))
 
# 针对结果再次find_all,从而获取所有的li标签信息
for ul in soup.find_all('ul'):
    print(ul.find_all('li'))
 
# 查找id为list-1的内容
print(soup.find_all(attrs={'id': 'list-1'}))
 
# 查找class为element的内容
print(soup.find_all(attrs={'class': 'element'}))
 
# 查找所有的text='foo'的文本
print(soup.find_all(text='foo'))

css选择器示例：

# -*- coding:utf-8 -*-
from bs4 import beautifulsoup
 
html = '''
<div class="panel">
    <div class="panel-heading">
        <h4>hello</h4>
    </div>
    <div class="panel-body">
        <ul class="list" id="list-1">
            <li class="element">foo</li>
            <li class="element">bar</li>
            <li class="element">jay</li>
        </ul>
        <ul class="list list-small" id="list-2">
            <li class="element">foo</li>
            <li class="element">bar</li>
        </ul>
    </div>
</div>
'''
 
soup = beautifulsoup(html, 'lxml')
 
# 获取class名为panel下panel-heading的内容
print(soup.select('.panel .panel-heading'))
 
# 获取class名为ul和li的内容
print(soup.select('ul li'))
 
# 获取class名为element，id为list-2的内容
print(soup.select('#list-2 .element'))
 
# 使用get_text()获取文本内容
for li in soup.select('li'):
    print(li.get_text())
 
# 获取属性的时候可以通过[属性名]或者attrs[属性名]
for ul in soup.select('ul'):
    print(ul['id'])
    # print(ul.attrs['id'])

您可能感兴趣的文章:

如对本文有疑问，请在下面进行留言讨论，广大热心网友会与你互动！！点击进行留言回复

python如何查看网页代码

用python查看网页代码的方法：1、使用“import”导入requests包import requests2、使用requests包的get()函数通过网页... [阅读全文]
Python如何用wx模块创建文本编辑器

用python的wx模块创建文本编辑器的方法：1、设置按钮的位置import wxapp = wx.app()win = wx.frame(none,title... [阅读全文]
python如何保存文本文件

python保存文本文件的方法：使用python内置的open()类可以打开文本文件，向文件里面写入数据可以用write()函数，写完之后，使用close()函... [阅读全文]
python如何编写win程序

python可以编写win程序。win程序的格式是exe，下面我们就来看一下使用python编写exe程序的方法。编写好python程序后py2exe模块即可将... [阅读全文]
Python替换NumPy数组中大于某个值的所有元素实例

我有一个2d(二维) numpy数组，并希望用255.0替换大于或等于阈值t的所有值。据我所知，最基础的方法是：shape = arr.shaperesult ... [阅读全文]
使用Numpy对特征中的异常值进行替换及条件替换方式

原始数据为excel文件，由传感器获得，通过pyhton xlrd模块读入，读入后为数组形式，由于其存在部分异常值和缺失值，所以便利用numpy对其中的异常值进... [阅读全文]
Python 实现将numpy中的nan和inf,nan替换成对应的均值

nan：not a numberinf：infinity;正无穷numpy中的nan和inf都是float类型t!=t 返回bool类型的数组(矩阵)np.co... [阅读全文]
给ubuntu18安装python3.7的详细教程

参考文章准备工作安装工具sudo apt updatesudo apt upgradesudo apt install gccsudo apt install ... [阅读全文]
python爬虫把url链接编码成gbk2312格式过程解析

1. 问题　　抓取某个网站，发现请求参数是乱码格式，这是点击 textview，发现请求参数如下图所示3. 那么=%b9%fa%ce%f1%d4%ba%b7%a... [阅读全文]
pyecharts在数据可视化中的应用详解

使用pyecharts进行数据可视化安装 pip install pyecharts也可以在pycharm软件里进行下载pyecharts库包。下载成功后进行查... [阅读全文]

网友评论


验证码：

Python BeautifulSoup 使用

2019年01月21日 | 移动技术网IT编程 | 我要评论

您可能感兴趣的文章:

相关文章:

网友评论