Monitoring signals worth keeping before enabling AI assistance
AI ใช้งานได้ดีเมื่อ telemetry มีคุณภาพ มีบริบท และมีผู้รับผิดชอบก่อนเริ่ม
AI begins with signal quality
AI assistance ใน cloud operations สามารถช่วยสรุป incident, เชื่อมโยงเหตุการณ์ และค้นหาข้อมูลด้วยภาษาธรรมชาติ แต่คำตอบยังขึ้นกับ signal และ context ที่ระบบเข้าถึงได้ หาก telemetry ขาดช่วง ชื่อ resource ไม่สื่อความหมาย หรือ alert ไม่มี owner ผลลัพธ์ที่ได้อาจรวดเร็วแต่ไม่ช่วยให้ตัดสินใจได้ดีขึ้น
จุดเริ่มต้นจึงไม่ใช่การเปิด AI ให้ครอบคลุมทุก subscription แต่เป็นการตรวจว่า monitoring ปัจจุบันอธิบาย health ของ workload ได้หรือไม่ และทีมรู้ว่าจะทำอะไรเมื่อ signal เปลี่ยน
Minimum useful signals
- Service health: availability, dependency failure และ platform event ที่มีผลต่อบริการ
- Workload health: request success, latency, error rate, queue depth และ saturation ที่สัมพันธ์กับ user journey
- Change context: deployment, configuration change, policy change และ privileged activity
- Security context: identity risk, suspicious access, control failure และ incident status
- Business context: service owner, criticality, environment, data classification, RTO และ RPO
รายการนี้ไม่จำเป็นต้องเก็บด้วยเครื่องมือเดียว แต่ควรมี identifier และ timestamp ที่ทำให้เชื่อมโยงกันได้ การตั้งชื่อ resource, tag และ environment อย่างสม่ำเสมอจึงมีผลต่อการวิเคราะห์มากกว่าที่เห็น
Context AI needs
Metrics ช่วยเห็นแนวโน้มและ threshold, logs ให้รายละเอียดของเหตุการณ์ และ traces ช่วยติดตามเส้นทางข้าม component ข้อมูลทั้งสามชนิดมีต้นทุนและวัตถุประสงค์ต่างกัน ควรกำหนด verbosity, category และ retention ตามสิ่งที่ต้องตรวจสอบ ไม่ใช่เปิดทุก category ไว้ตลอดเวลา
AI ยังต้องรู้ว่า signal ใดสัมพันธ์กับบริการใด ใครเป็นเจ้าของ และการเปลี่ยนแปลงใดเกิดขึ้นก่อนหน้า Health model หรือ service map ที่เชื่อม infrastructure กับ business function ช่วยลดการสรุปจาก log แบบแยกส่วน
Operational steps
- เลือกหนึ่ง workload และนิยาม user journey กับ health indicator ที่สำคัญ
- ทำ inventory ของ metrics, logs, traces, alerts และ change records ที่มีอยู่
- ปิด signal ที่ซ้ำหรือไม่มี action และเติมช่องว่างที่ทำให้วิเคราะห์ incident ไม่ครบ
- เชื่อม service owner, environment, criticality และ deployment context เข้ากับ telemetry
- ทดลอง AI แบบ read-only กับ incident ย้อนหลัง แล้วเปรียบเทียบคำตอบกับ evidence จริง
- เพิ่ม action ทีละระดับ โดยกำหนด approval, RBAC, audit trail และวิธีหยุด automation ก่อนใช้กับ production
Readiness check
ก่อนให้ AI เข้ามาช่วย workflow ควรตอบคำถามเหล่านี้ให้ได้
- Alert นี้สะท้อนผลต่อผู้ใช้หรือเป็นเพียงความผิดปกติของ resource
- มี runbook หรือ decision boundary ที่ใช้ตรวจคำแนะนำหรือไม่
- ข้อมูลที่ AI เข้าถึงมี secret, personal data หรือขอบเขตที่ห้ามส่งต่อหรือไม่
- คำตอบและ action ถูกบันทึกเพื่อตรวจสอบย้อนหลังได้หรือไม่
- ใครเป็นผู้อนุมัติ remediation และจะหยุด automation ได้อย่างไร
Human boundary
งานที่มีผลสูง เช่น ปิด resource, เปลี่ยน network route, revoke access หรือแก้ production configuration ควรมี human review และใช้สิทธิ์ตาม role ที่กำหนด AI สามารถเสนอขั้นตอนหรือรวบรวม evidence ได้ แต่ operational ownership ยังอยู่กับทีมที่รับผิดชอบระบบ
AI ไม่ได้แก้ monitoring ที่ไม่มี owner แต่ทำให้ข้อจำกัดของ monitoring นั้นปรากฏเร็วขึ้น