核心交换机一坏全楼断网?华为CSS集群两台虚拟成一台坏一台照跑
企业网做到核心层,几乎都会上两台交换机做冗余。但传统做法是两台各自独立跑 VRRP 加双链路,配置要在两台上各做一遍、链路聚合只能框内捆绑、上下行设备还得分别跟两台核心对接,复杂度翻倍。华为交换机还有一条路——集群交换机系统 CSS,把两台支持集群特性的交换机在逻辑上组合成一台交换设备:对全网呈现一个 IP 地址一个 MAC 地址,运维只需登录任意一台成员就能管理整个集群,成员之间冗余备份,再配合跨设备的链路聚合把可靠性做到链路级。
集群成员之间怎么连?主流的是集群卡方式——通过主控板或交换网板上的集群卡及集群线缆连接,不占普通业务端口,配置简单、稳定性高、时延小。本文以 S9706 为例,核心层 SwitchA 与 SwitchB 采取集群卡集群组网,SwitchA 为主交换机,SwitchB 为备交换机;汇聚层交换机通过 Eth-Trunk 双归接入集群系统,集群系统再通过 Eth-Trunk 接入上行网络。整套做下来分五步。
配置步骤
第一步,连集群线缆。按照设备款型及集群卡型号选择相应的连线图连接,连线规则的核心就一条——相同编号及相同颜色的接口全部对接(如左框蓝1必须连接右框蓝1)。以 VSTSA 集群卡为例每块卡有 4 个接口,这种卡的连线最多只允许一根线缆故障,所以务必按图连满。
第二步,配集群 ID 与优先级。两台设备的集群 ID 必须不同,优先级用来决定谁当主交换机。SwitchA 保持缺省集群 ID 1,优先级配成 100:
<HUAWEI> system-view
[HUAWEI] sysname SwitchA
[SwitchA] set css priority 100
SwitchB 的集群 ID 配成 2,优先级 10:
<HUAWEI> system-view
[HUAWEI] sysname SwitchB
[SwitchB] set css id 2
[SwitchB] set css priority 10
配完先自查一遍,确认保存的配置与预期一致:
[SwitchA] display css status saved
Current Id Saved Id CSS Enable CSS Mode Priority Master force
------------------------------------------------------------------------------
1 1 Off CSS card 100 Off
[SwitchB] display css status saved
Current Id Saved Id CSS Enable CSS Mode Priority Master force
------------------------------------------------------------------------------
1 2 Off CSS card 10 Off
第三步,使能集群功能。注意顺序——先使能 SwitchA 并重启,再使能 SwitchB,以保证 SwitchA 成为主交换机。两台都会提示重启后才生效:
[SwitchA] css enable
Warning: The CSS configuration will take effect only after the system is rebooted. T
he next CSS mode is CSS card. Reboot now? [Y/N]:y
[SwitchB] css enable
Warning: The CSS configuration will take effect only after the system is rebooted. T
he next CSS mode is CSS card. Reboot now? [Y/N]:y
第四步,确认集群组建成功。先看灯——SwitchA 集群卡上 MASTER 灯常亮表示它是主交换机,SwitchB 集群卡上 MASTER 灯常灭表示是备交换机。再通过 Console 口登录集群系统执行 display device,输出中能看到 Chassis 1(Master Switch)与 Chassis 2(Standby Switch)两框的单板状态,说明集群建立完成。最后看集群链路:
<SwitchA> display css channel
Chassis 1 || Chassis 2
================================================================================
Num [SRUC HG] [VS08 Port(Status)] || [VS08 Port(Status)] [SRUC HG]
1 1/7 0/12 -- 1/7/0/1(UP 10G) ---||--- 2/7/0/1(UP 10G) -- 2/7 0/12
2 1/7 0/16 -- 1/7/0/2(UP 10G) ---||--- 2/7/0/2(UP 10G) -- 2/7 0/16
3 1/7 0/13 -- 1/7/0/3(UP 10G) ---||--- 2/7/0/3(UP 10G) -- 2/7 0/13
4 1/7 0/17 -- 1/7/0/4(UP 10G) ---||--- 2/7/0/4(UP 10G) -- 2/7 0/17
5 1/7 0/14 -- 1/7/0/5(UP 10G) ---||--- 2/8/0/5(UP 10G) -- 2/8 0/14
6 1/7 0/18 -- 1/7/0/6(UP 10G) ---||--- 2/8/0/6(UP 10G) -- 2/8 0/18
7 1/7 0/15 -- 1/7/0/7(UP 10G) ---||--- 2/8/0/7(UP 10G) -- 2/8 0/15
8 1/7 0/19 -- 1/7/0/8(UP 10G) ---||--- 2/8/0/8(UP 10G) -- 2/8 0/19
9 1/8 0/12 -- 1/8/0/1(UP 10G) ---||--- 2/8/0/1(UP 10G) -- 2/8 0/12
10 1/8 0/16 -- 1/8/0/2(UP 10G) ---||--- 2/8/0/2(UP 10G) -- 2/8 0/16
11 1/8 0/13 -- 1/8/0/3(UP 10G) ---||--- 2/8/0/3(UP 10G) -- 2/8 0/13
12 1/8 0/17 -- 1/8/0/4(UP 10G) ---||--- 2/8/0/4(UP 10G) -- 2/8 0/17
13 1/8 0/14 -- 1/8/0/5(UP 10G) ---||--- 2/7/0/5(UP 10G) -- 2/7 0/14
14 1/8 0/18 -- 1/8/0/6(UP 10G) ---||--- 2/7/0/6(UP 10G) -- 2/7 0/18
15 1/8 0/15 -- 1/8/0/7(UP 10G) ---||--- 2/7/0/7(UP 10G) -- 2/7 0/15
16 1/8 0/19 -- 1/8/0/8(UP 10G) ---||--- 2/7/0/8(UP 10G) -- 2/7 0/19
所有集群链路均为 UP,集群组建完全成功。顺手把集群系统改名,方便后续识别,并配置上下行 Eth-Trunk——上行 Eth-Trunk 10、对 SwitchC 的下行 Eth-Trunk 20、对 SwitchD 的下行 Eth-Trunk 30,成员口分布在两台成员交换机上,这就是跨设备聚合:
<SwitchA> system-view
[SwitchA] sysname CSS //给集群系统重新命名
[CSS] interface eth-trunk 10
[CSS-Eth-Trunk10] quit
[CSS] interface gigabitethernet 1/1/0/4
[CSS-GigabitEthernet1/1/0/4] eth-trunk 10
[CSS-GigabitEthernet1/1/0/4] quit
[CSS] interface gigabitethernet 2/1/0/4
[CSS-GigabitEthernet2/1/0/4] eth-trunk 10
[CSS-GigabitEthernet2/1/0/4] quit
[CSS] interface eth-trunk 20
[CSS-Eth-Trunk20] quit
[CSS] interface gigabitethernet 1/1/0/3
[CSS-GigabitEthernet1/1/0/3] eth-trunk 20
[CSS-GigabitEthernet1/1/0/3] quit
[CSS] interface gigabitethernet 2/1/0/5
[CSS-GigabitEthernet2/1/0/5] eth-trunk 20
[CSS-GigabitEthernet2/1/0/5] quit
[CSS] interface eth-trunk 30
[CSS-Eth-Trunk30] quit
[CSS] interface gigabitethernet 1/1/0/5
[CSS-GigabitEthernet1/1/0/5] eth-trunk 30
[CSS-GigabitEthernet1/1/0/5] quit
[CSS] interface gigabitethernet 2/1/0/3
[CSS-GigabitEthernet2/1/0/3] eth-trunk 30
[CSS-GigabitEthernet2/1/0/3] return
对端设备按同样思路建 Eth-Trunk 与集群对接,以上行 SwitchE 为例(SwitchC、SwitchD 同理):
<HUAWEI> system-view
[HUAWEI] sysname SwitchE
[SwitchE] interface eth-trunk 10
[SwitchE-Eth-Trunk10] quit
[SwitchE] interface gigabitethernet 1/0/1
[SwitchE-GigabitEthernet1/0/1] eth-trunk 10
[SwitchE-GigabitEthernet1/0/1] quit
[SwitchE] interface gigabitethernet 1/0/2
[SwitchE-GigabitEthernet1/0/2] eth-trunk 10
[SwitchE-GigabitEthernet1/0/2] quit
检查聚合口状态,成员口两个都在 trunk 里且 operate up:
<CSS> display trunkmembership eth-trunk 10
Trunk ID: 10
Used status: VALID
TYPE: ethernet
Working Mode : Normal
Number Of Ports in Trunk = 2
Number Of Up Ports in Trunk = 2
Operate status: up
Interface GigabitEthernet1/1/0/4, valid, operate up, weight=1
Interface GigabitEthernet2/1/0/4, valid, operate up, weight=1
第五步,配多主检测。很多机房会漏这步——集群成员共用同一个 IP 和 MAC,集群线缆全断发生分裂时,网络里会出现两台"同一身份"的设备引发地址冲突。MAD 多主检测就是防这个的,这里用代理方式、由 SwitchC 兼任代理设备,不额外占用端口。在集群系统的下行 Eth-Trunk 20 上开启:
<CSS> system-view
[CSS] interface eth-trunk 20
[CSS-Eth-Trunk20] mad detect mode relay //V200R002C00及之前的版本命令行格式为dual-active detect mode relay
[CSS-Eth-Trunk20] quit
[CSS] quit
再在 SwitchC 上开启代理功能:
[SwitchC] interface eth-trunk 20
[SwitchC-Eth-Trunk20] mad relay //V200R002C00及之前的版本命令行格式为dual-active relay
[SwitchC-Eth-Trunk20] quit
[SwitchC] quit
验证两边都生效:
<CSS> display mad //V200R002C00及之前的版本命令行格式为display dual-active
Current MAD domain: 0
MAD direct detection enabled: NO
MAD relay detection enabled: YES
<SwitchC> display mad proxy //V200R002C00及之前的版本命令行格式为display dual-active proxy
Mad relay interfaces configured:
Eth-Trunk20
集群系统与 SwitchC 的最终配置文件如下(SwitchD、SwitchE 结构相同,分别对应 Eth-Trunk 30 与 Eth-Trunk 10):
#
sysname CSS
#
interface Eth-Trunk10
#
interface Eth-Trunk20
mad detect mode relay
#
interface Eth-Trunk30
#
interface GigabitEthernet1/1/0/3
eth-trunk 20
#
interface GigabitEthernet1/1/0/4
eth-trunk 10
#
interface GigabitEthernet1/1/0/5
eth-trunk 30
#
interface GigabitEthernet2/1/0/3
eth-trunk 30
#
interface GigabitEthernet2/1/0/4
eth-trunk 10
#
interface GigabitEthernet2/1/0/5
eth-trunk 20
#
return
#
sysname SwitchC
#
interface Eth-Trunk20
mad relay
#
interface GigabitEthernet1/0/1
eth-trunk 20
#
interface GigabitEthernet1/0/2
eth-trunk 20
#
return
有两个运维细节值得记下:集群重启或更换主控板后集群 MAC 可能变化,若业务依赖固定 MAC,可提前用 set css system-mac 把它固定成某台成员的 MAC;另外直连和代理两种 MAD 检测方式互斥,配了集群 Eth-Trunk 的场景建议就选代理方式。贵州诚鑫致达科技在机房核心层改造中做过多次双机集群落地,分裂检测这类补丁配置都是验收清单里的必检项。你机房的核心层是堆叠、集群还是 VRRP 双机?评论区聊聊各自的运维感受。