scray中设置动态ip - 代码天地

scray中设置动态ip

编程语言 2018-11-03 05:56:50 阅读次数: 0

在scray写入一个脚本文件

思路：

1、建立def crawl_ips():方法获取爬取免费的西刺ip。

2、将获取的ip存放在数据库中，对其进行判断分析，剔除无效ip

3、在middleware文件进行获取保存好的有效ip

爬取ip代理脚本
import requests
from scrapy.selector import Selector
import MySQLdb
conn=MySQLdb.connect(host="192.168.0.104",user="root",passwd="spiderliu",db="article_spider",charest="utf8")
cursor=conn.cursor()

#设置ip代理,可以使用类似item方式进行
def crawl_ips():
#爬取西刺免费代理
    headers={"UserAgent":"Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/63.0.3239.132 Safari/537.36"}
    for i in range(3473):
    #通过requests获取整个页面的信息
        re=requests.get("http://www.xicidaili.com/nn/{0}".format(i),headers=headers)
    #通过selector获取页面
    selector=Selector(text=re.text)
    #获取所有tr
    all_trs=selector.css("#ip_list tr")
    ip_list=[]
    for tr in all_trs[1:]:
        # 获取速度
        speed_str=tr.css(".bar::attr(title)").extract()[0]
        if speed_str:
            speed=float(speed_str.split("秒")[0])
        all_texts=tr.css("td::text").extract()
        ip=all_texts[0]
        port=all_texts[1]
        proxy_type=all_texts[5]
        ip_list.append((ip,port,proxy_type,speed))
    for ip_info in ip_list:
        cursor.execute(
            "insert proxy_ip(ip,port,proxy_type,speed) VALUES('{0}','{1}',{2},'http or https')".format(
                ip_info[0], ip_info[1], ip_info[3]
            )
        )
        conn.commit()

class GetIP(object):
    def delete_ip(self,ip):
        #从数据库中删除无效的ip
        delete_sql="""
            delete from proxy_ip where ip='{0}'
        """.format(ip)
        cursor.execute(delete_sql)
        cursor.commit()
        return True
    #对于获取的ip进行判断
    def judge_ip(self,ip,port):
        http_url="www.baidu.com"
        proxy_url="https://{0}:{1}".format(ip,port)
        try:
            proxy_dict={
                "https":proxy_url
            }
            response=requests.get(http_url,proxies=proxy_dict)
        except Exception as e:
            print("无效ip和端口")
            self.delete_ip()
            return False
        else:
            code=response.status_code
            if code>=200 and code<300:
                print("有效的ip")
                return True
            else:
                print("无效")
                self.delete_ip()
                return False

    #从数据库中随机获取一个ip
    def get_random_ip(self):
        random_sql="""
              SELECT ip,prot FROM proxy_ip
            ORDER BY RAND()
            LIMIT 1
        """
        result=cursor.execute(random_sql)
        for ip_info in cursor.fetchall():
            ip=ip_info[0]
            port=ip_info[1]
            judge_re=self.judge_ip(ip,port)
            if judge_re:
                return "http://{0}:{1}".format(ip,port)
            else:
                return self.get_random_ip()

#print(crawl_ips())
if name=="main":
get_ip=GetIP()
get_ip.get_random_ip()

middlewares文件中获取脚本中的ip
class RandomProxyMiddleware(object):
    #动态ip设置
    def process_request(self,request,spider):
        get_ip=GetIP()
        request.meta["proxy"]=get_ip.get_random_ip()

猜你喜欢

转载自blog.csdn.net/lx516109011/article/details/83572910

scray中设置动态ip

scray中的Request 不执行回调解决办法

如何设置动态IP上网？

电脑怎么设置动态IP

Linux—— 静态IP跟动态IP设置

linux_静态IP与动态IP设置

Python爬虫实例九州动态IP使用HTTP的urllib2中的ProxyHandler设置。

设置电脑动态换IP的方法

设置MAC地址和动态IP

电脑如何设置动态IP上网

电脑怎么设置修改成动态IP

scrapy中使用动态转发IP设置

简单几步设置电脑动态换ip

路由器动态ip设置

一键把动态IP自动设置为静态IP

Linux下设置静态IP和获取动态IP的方法

Linux eth0设置静态ip和动态ip

Ubuntu将动态ip设置为静态ip（NAT方式）

请求中设置代理IP

CentOS中设置静态IP

linux中设置静态ip

在selenium中设置代理ip

docker中设置固定IP

在Centos中设置固定ip

linux中设置固定ip

linux中静态ip的设置

IP地址静态设置和动态设置区别

scrapy中自定义下载中间件设置动态User-Agent和代理ip

middlewares中自己设置UA的切换，随机更换IP地址，以及对#selenium 动态页面的爬取

动态设置标签中的内容

今日推荐

Linus “吃狗粮”最积极！

开源日报 | Winamp播放器即将开源；生成式AI之战升级第二轮；Linus“吃狗粮”最积极；AI进入泡沫前期；吴泳铭为阿里云带来了什么？

NetBSD 禁止提交由 AI 生成的代码

Apache Doris 2.0.10 版本正式发布！

开源日报 | 大模型开战；大模型独角兽被曝卖身；周鸿祎建议谷歌开源所有产品；最大开源AI社区提供1000万美元共享GPU

开源日报 | Chrome内置Gemini的意义不在于Gemini；中国AI追随之路的五大误区；ECharts创始人“下海”养鱼；谷歌I/O开发者大会什么都有，只是没有惊喜

微软回应中国区AI团队“打包赴美”传闻

周排行

LogN级别的区间查询算法(线段树), 你学会了吗

数论概论(英文版.第4版)

idea 更新后和新的直接安装前，都需要配置 idea64.exe.vmoptions 后再使用

CANOpen系列教程04_CAN总线波特率、位时序、帧类型及格式说明

Java序列化基础

java排序算法整理

异常：org.apache.ibatis.reflection.ReflectionException

（算法练习）——二路归并排序

go 闭包函数

好程序员web前端技术分享媒体查询

每日归档

更多

2024-05-21(8)

2024-05-20(36)

2024-05-19(0)

2024-05-18(4)

2024-05-17(34)

2024-05-16(6)

2024-05-15(24)

2024-05-14(0)

2024-05-13(18)

2024-05-12(0)