核心交换机一坏全楼断网?华为CSS集群两台虚拟成一台坏一台照跑

企业网做到核心层,几乎都会上两台交换机做冗余。但传统做法是两台各自独立跑 VRRP 加双链路,配置要在两台上各做一遍、链路聚合只能框内捆绑、上下行设备还得分别跟两台核心对接,复杂度翻倍。华为交换机还有一条路——集群交换机系统 CSS,把两台支持集群特性的交换机在逻辑上组合成一台交换设备:对全网呈现一个 IP 地址一个 MAC 地址,运维只需登录任意一台成员就能管理整个集群,成员之间冗余备份,再配合跨设备的链路聚合把可靠性做到链路级。

集群成员之间怎么连?主流的是集群卡方式——通过主控板或交换网板上的集群卡及集群线缆连接,不占普通业务端口,配置简单、稳定性高、时延小。本文以 S9706 为例,核心层 SwitchA 与 SwitchB 采取集群卡集群组网,SwitchA 为主交换机,SwitchB 为备交换机;汇聚层交换机通过 Eth-Trunk 双归接入集群系统,集群系统再通过 Eth-Trunk 接入上行网络。整套做下来分五步。

配置步骤

第一步,连集群线缆。按照设备款型及集群卡型号选择相应的连线图连接,连线规则的核心就一条——相同编号及相同颜色的接口全部对接(如左框蓝1必须连接右框蓝1)。以 VSTSA 集群卡为例每块卡有 4 个接口,这种卡的连线最多只允许一根线缆故障,所以务必按图连满。

第二步,配集群 ID 与优先级。两台设备的集群 ID 必须不同,优先级用来决定谁当主交换机。SwitchA 保持缺省集群 ID 1,优先级配成 100:

<HUAWEI> system-view
[HUAWEI] sysname SwitchA
[SwitchA] set css priority 100

SwitchB 的集群 ID 配成 2,优先级 10:

<HUAWEI> system-view
[HUAWEI] sysname SwitchB
[SwitchB] set css id 2
[SwitchB] set css priority 10

配完先自查一遍,确认保存的配置与预期一致:

[SwitchA] display css status saved
Current Id   Saved Id     CSS Enable   CSS Mode    Priority    Master force     
------------------------------------------------------------------------------  
1            1            Off          CSS card    100         Off             
[SwitchB] display css status saved
Current Id   Saved Id     CSS Enable   CSS Mode    Priority    Master force     
------------------------------------------------------------------------------  
1            2            Off          CSS card    10          Off              

第三步,使能集群功能。注意顺序——先使能 SwitchA 并重启,再使能 SwitchB,以保证 SwitchA 成为主交换机。两台都会提示重启后才生效:

[SwitchA] css enable
Warning: The CSS configuration will take effect only after the system is rebooted. T
he next CSS mode is CSS card. Reboot now? [Y/N]:y
[SwitchB] css enable
Warning: The CSS configuration will take effect only after the system is rebooted. T
he next CSS mode is CSS card. Reboot now? [Y/N]:y

第四步,确认集群组建成功。先看灯——SwitchA 集群卡上 MASTER 灯常亮表示它是主交换机,SwitchB 集群卡上 MASTER 灯常灭表示是备交换机。再通过 Console 口登录集群系统执行 display device,输出中能看到 Chassis 1(Master Switch)与 Chassis 2(Standby Switch)两框的单板状态,说明集群建立完成。最后看集群链路:

<SwitchA> display css channel
                Chassis 1               ||               Chassis 2              
================================================================================
Num [SRUC HG]    [VS08 Port(Status)]    ||    [VS08 Port(Status)]    [SRUC HG]  
1   1/7  0/12 -- 1/7/0/1(UP 10G)     ---||--- 2/7/0/1(UP 10G)     -- 2/7  0/12  
2   1/7  0/16 -- 1/7/0/2(UP 10G)     ---||--- 2/7/0/2(UP 10G)     -- 2/7  0/16  
3   1/7  0/13 -- 1/7/0/3(UP 10G)     ---||--- 2/7/0/3(UP 10G)     -- 2/7  0/13  
4   1/7  0/17 -- 1/7/0/4(UP 10G)     ---||--- 2/7/0/4(UP 10G)     -- 2/7  0/17  
5   1/7  0/14 -- 1/7/0/5(UP 10G)     ---||--- 2/8/0/5(UP 10G)     -- 2/8  0/14  
6   1/7  0/18 -- 1/7/0/6(UP 10G)     ---||--- 2/8/0/6(UP 10G)     -- 2/8  0/18  
7   1/7  0/15 -- 1/7/0/7(UP 10G)     ---||--- 2/8/0/7(UP 10G)     -- 2/8  0/15  
8   1/7  0/19 -- 1/7/0/8(UP 10G)     ---||--- 2/8/0/8(UP 10G)     -- 2/8  0/19  
9   1/8  0/12 -- 1/8/0/1(UP 10G)     ---||--- 2/8/0/1(UP 10G)     -- 2/8  0/12  
10  1/8  0/16 -- 1/8/0/2(UP 10G)     ---||--- 2/8/0/2(UP 10G)     -- 2/8  0/16  
11  1/8  0/13 -- 1/8/0/3(UP 10G)     ---||--- 2/8/0/3(UP 10G)     -- 2/8  0/13  
12  1/8  0/17 -- 1/8/0/4(UP 10G)     ---||--- 2/8/0/4(UP 10G)     -- 2/8  0/17  
13  1/8  0/14 -- 1/8/0/5(UP 10G)     ---||--- 2/7/0/5(UP 10G)     -- 2/7  0/14  
14  1/8  0/18 -- 1/8/0/6(UP 10G)     ---||--- 2/7/0/6(UP 10G)     -- 2/7  0/18  
15  1/8  0/15 -- 1/8/0/7(UP 10G)     ---||--- 2/7/0/7(UP 10G)     -- 2/7  0/15  
16  1/8  0/19 -- 1/8/0/8(UP 10G)     ---||--- 2/7/0/8(UP 10G)     -- 2/7  0/19  

所有集群链路均为 UP,集群组建完全成功。顺手把集群系统改名,方便后续识别,并配置上下行 Eth-Trunk——上行 Eth-Trunk 10、对 SwitchC 的下行 Eth-Trunk 20、对 SwitchD 的下行 Eth-Trunk 30,成员口分布在两台成员交换机上,这就是跨设备聚合:

<SwitchA> system-view
[SwitchA] sysname CSS              //给集群系统重新命名
[CSS] interface eth-trunk 10
[CSS-Eth-Trunk10] quit
[CSS] interface gigabitethernet 1/1/0/4
[CSS-GigabitEthernet1/1/0/4] eth-trunk 10
[CSS-GigabitEthernet1/1/0/4] quit
[CSS] interface gigabitethernet 2/1/0/4
[CSS-GigabitEthernet2/1/0/4] eth-trunk 10
[CSS-GigabitEthernet2/1/0/4] quit
[CSS] interface eth-trunk 20
[CSS-Eth-Trunk20] quit
[CSS] interface gigabitethernet 1/1/0/3
[CSS-GigabitEthernet1/1/0/3] eth-trunk 20
[CSS-GigabitEthernet1/1/0/3] quit
[CSS] interface gigabitethernet 2/1/0/5
[CSS-GigabitEthernet2/1/0/5] eth-trunk 20
[CSS-GigabitEthernet2/1/0/5] quit
[CSS] interface eth-trunk 30
[CSS-Eth-Trunk30] quit
[CSS] interface gigabitethernet 1/1/0/5
[CSS-GigabitEthernet1/1/0/5] eth-trunk 30
[CSS-GigabitEthernet1/1/0/5] quit
[CSS] interface gigabitethernet 2/1/0/3
[CSS-GigabitEthernet2/1/0/3] eth-trunk 30
[CSS-GigabitEthernet2/1/0/3] return

对端设备按同样思路建 Eth-Trunk 与集群对接,以上行 SwitchE 为例(SwitchC、SwitchD 同理):

<HUAWEI> system-view
[HUAWEI] sysname SwitchE
[SwitchE] interface eth-trunk 10
[SwitchE-Eth-Trunk10] quit
[SwitchE] interface gigabitethernet 1/0/1
[SwitchE-GigabitEthernet1/0/1] eth-trunk 10
[SwitchE-GigabitEthernet1/0/1] quit
[SwitchE] interface gigabitethernet 1/0/2
[SwitchE-GigabitEthernet1/0/2] eth-trunk 10
[SwitchE-GigabitEthernet1/0/2] quit

检查聚合口状态,成员口两个都在 trunk 里且 operate up:

<CSS> display trunkmembership eth-trunk 10
Trunk ID: 10
Used status: VALID
TYPE: ethernet
Working Mode : Normal
Number Of Ports in Trunk = 2
Number Of Up Ports in Trunk = 2
Operate status: up
Interface GigabitEthernet1/1/0/4, valid, operate up, weight=1
Interface GigabitEthernet2/1/0/4, valid, operate up, weight=1

第五步,配多主检测。很多机房会漏这步——集群成员共用同一个 IP 和 MAC,集群线缆全断发生分裂时,网络里会出现两台"同一身份"的设备引发地址冲突。MAD 多主检测就是防这个的,这里用代理方式、由 SwitchC 兼任代理设备,不额外占用端口。在集群系统的下行 Eth-Trunk 20 上开启:

<CSS> system-view
[CSS] interface eth-trunk 20
[CSS-Eth-Trunk20] mad detect mode relay           //V200R002C00及之前的版本命令行格式为dual-active detect mode relay
[CSS-Eth-Trunk20] quit
[CSS] quit

再在 SwitchC 上开启代理功能:

[SwitchC] interface eth-trunk 20
[SwitchC-Eth-Trunk20] mad relay                    //V200R002C00及之前的版本命令行格式为dual-active relay
[SwitchC-Eth-Trunk20] quit
[SwitchC] quit

验证两边都生效:

<CSS> display mad                                 //V200R002C00及之前的版本命令行格式为display dual-active 
Current MAD domain: 0  
MAD direct detection enabled: NO
MAD relay detection enabled: YES
<SwitchC> display mad proxy                      //V200R002C00及之前的版本命令行格式为display dual-active proxy
Mad relay interfaces configured:
 Eth-Trunk20

集群系统与 SwitchC 的最终配置文件如下(SwitchD、SwitchE 结构相同,分别对应 Eth-Trunk 30 与 Eth-Trunk 10):

#
sysname CSS
#
interface Eth-Trunk10
#
interface Eth-Trunk20
 mad detect mode relay
#
interface Eth-Trunk30
#
interface GigabitEthernet1/1/0/3
 eth-trunk 20
#
interface GigabitEthernet1/1/0/4
 eth-trunk 10
#
interface GigabitEthernet1/1/0/5
 eth-trunk 30
#
interface GigabitEthernet2/1/0/3
 eth-trunk 30
#
interface GigabitEthernet2/1/0/4
 eth-trunk 10
#
interface GigabitEthernet2/1/0/5
 eth-trunk 20
#
return
#
sysname SwitchC
#
interface Eth-Trunk20
 mad relay
#
interface GigabitEthernet1/0/1
 eth-trunk 20
#
interface GigabitEthernet1/0/2
 eth-trunk 20
#
return

有两个运维细节值得记下:集群重启或更换主控板后集群 MAC 可能变化,若业务依赖固定 MAC,可提前用 set css system-mac 把它固定成某台成员的 MAC;另外直连和代理两种 MAD 检测方式互斥,配了集群 Eth-Trunk 的场景建议就选代理方式。贵州诚鑫致达科技在机房核心层改造中做过多次双机集群落地,分裂检测这类补丁配置都是验收清单里的必检项。你机房的核心层是堆叠、集群还是 VRRP 双机?评论区聊聊各自的运维感受。