import requests from lxml import etree url = 'https://jn.lianjia.com/zufang/' headers = { 'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/93.0.4577.63 Safari/537.36' } response = requests.get(url, headers=headers) html = response.text selector = etree.HTML(html) total_page_list = selector.xpath('//div[@class="content__pg"]/div/@data-totalpage') if len(total_page_list) > 0: total_page = total_page_list[0] print('总页数为:', total_page) else: print('未匹配到结果')报错
时间: 2024-04-01 11:34:46 浏览: 77
from lxml import etree
这段代码没有报错,但是你忘记了循环爬取每一页的租房信息。你可以将代码改为以下形式:
```
import requests
from lxml import etree
url = 'https://jn.lianjia.com/zufang/'
headers = {
'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/93.0.4577.63 Safari/537.36'
}
response = requests.get(url, headers=headers)
html = response.text
selector = etree.HTML(html)
total_page_list = selector.xpath('//div[@class="content__pg"]/div/@data-totalpage')
if len(total_page_list) > 0:
total_page = total_page_list[0]
print('总页数为:', total_page)
# 循环爬取每一页的租房信息
for num in range(1, int(total_page) + 1):
url = 'https://jn.lianjia.com/zufang/pg{}/#contentList'.format(num)
# 发送请求并处理响应
# ...
else:
print('未匹配到结果')
```
这段代码首先发送一个请求到济南市链家网页,然后使用lxml库中的etree模块将网页转化为可操作的对象,接着使用xpath定位到包含总页数的元素,并进行判断。如果匹配成功,就输出总页数,并循环爬取每一页的租房信息。在循环中,使用.format()方法动态生成每一页的URL,并发送请求并处理响应。如果未匹配到结果,则输出相关信息。
阅读全文