python获取整个网页源码的方法 - CSDN文库

162 浏览量更新于2023-05-04 收藏 118KB PDF 举报

身份认证购VIP最低享 7 折!

领优惠券(最高得80元）

资源详情

资源推荐

python获取整个网页源码的方法获取整个网页源码的方法

1、、Python中获取整个页面的代码：中获取整个页面的代码：

import requests

res = requests.get('https://blog.csdn.net/yirexiao/article/details/79092355')

res.encoding = 'utf-8'

print(res.text)

2、运行结果、运行结果

实例扩展：实例扩展：

from bs4 import BeautifulSoup

import time,re,urllib2

t=time.time()

websiteurls={}

def scanpage(url):

websiteurl=url

t=time.time()

n=0

html=urllib2.urlopen(websiteurl).read()

soup=BeautifulSoup(html)

pageurls=[] Upageurls={}

pageurls=soup.find_all("a",href=True)

for links in pageurls:

if websiteurl in links.get("href") and links.get("href") not in Upageurls and links.get("href") not in websiteurls:

Upageurls[links.get("href")]=0

for links in Upageurls.keys():

try:

urllib2.urlopen(links).getcode()

except:

print "connect failed"

else:

t2=time.time()

Upageurls[links]=urllib2.urlopen(links).getcode()

print n,

print links,

print Upageurls[links] t1=time.time()

print t1-t2

n+=1

print ("total is "+repr(n)+" links")

print time.time()-t

scanpage(http://news.163.com/)

您可能感兴趣的文章您可能感兴趣的文章:Python爬虫获取页面所有URL链接过程详解python3+selenium获取页面加载的所有静态资源文件链接操作python

xpath获取页面注释的方法python 获取页面表格数据存放到csv中的方法Python get获取页面cookie代码实例Python基于lxml模块解析html

获取页面内所有叶子节点xpath路径功能示例

本内容试读结束，登录后可阅读更多

下载后可阅读完整内容，剩余0页未读，立即下载

weixin_38631042

粉丝: 4
资源: 927

会员权益专享

图片转文字

全年可省5，000元立即开通

最新资源

资源上传下载、课程学习等过程中有任何疑问或建议，欢迎提出宝贵意见哦~我们会及时处理！点击此处反馈