DevOps&SRE Library
описание
Библиотека статей по теме DevOps и SRE. Реклама: @ostinostin Контент: @mxssl РКН: https://www.gosuslugi.ru/snet/67704b536aa9672b963777b3
19 729
подписчиков
Охват к подписчикам
11,2%
ERR
Реакции к просмотрам
0,00%
0 на 27 постов
Пересылки к просмотрам
0,84%
506
Постов в день
2,6
всего 27
Где отзываются чаще
доля реакций к просмотрам- 09:03без подписи0,00%
- 17:01без подписи0,00%
- 15 авг.без подписи0,00%
- 14 авг.Migrating from Slurm to Kubernetes Moving from Slurm to Kubernetes doesn't have to mean losing the workflow you know. Here's how SkyPilot brings Slurm-like simplicity to K8s. https://blog.skypilot.co/slurm-to-k8s-migration0,00%
- 14 авг.The Hybrid Cloud Platform Illusion: Why Your On-Prem and Cloud Are Still Strangers I've spent the last few months working on what should have been a solved problem: letting applications running in our on-premises Kubernetes clusters access Google Cloud services. What I found instead was an industry-wide workaround culture built on security anti-patterns, and a surprisingly elegant solution hiding in plain sight. https://medium.com/@shkatara/the-hybrid-cloud-platform-illusion-why-your-on-prem-and-cloud-are-still-strangers-234a90ad89f10,00%
- 13 авг.VibeOps: A Secure read-only setup for AI-Assisted Kubernetes Debugging There is a lot of noise right now about letting AI "fix" your infrastructure. When production is acting up, you need to maintain a complete mental model of the system. If you let the AI be the driving force, you lose the overview. https://simon-frey.com/blog/vibeops-kubernetes0,00%
- 13 авг.Stop Manually Generating Kubeconfigs: Meet KubeUser KubeUser is a kubernetes-native operator that turns user management into a declarative code experience. No more manual certificate juggling — just apply a YAML file, and the operator handles the rest. https://medium.com/@yahya.muhaned/stop-manually-generating-kubeconfigs-meet-kubeuser-2f3ca87b027a0,00%
- 12 авг.You Don't Have a GIL Problem — You Have a CPU Problem This article documents a real production investigation into latency variance in Python-based microservices running on Kubernetes, revealing how CPU throttling amplifies GIL contention into unpredictable response time spikes. https://medium.com/@prashant_pathak/you-dont-have-a-gil-problem-you-have-a-cpu-problem-24deeadfea4a0,00%
- 12 авг.Разработчики получают инфраструктуру самостоятельно. DevOps — перестают выполнять однотипные запросы. 18 августа на бесплатном онлайн-вебинаре Orion soft покажет, как работает новая IDP-функциональность HyperDrive: self-service, GitOps, политики безопасности и управление инфраструктурой через Model Context Protocol. Реклама. ООО "Орион", ИНН: ИНН 9704113582, erid: 2Vtzqvqgags0,00%
- 12 авг.The feedback loops behind Kubernetes For the last decade, Kubernetes has been the backdrop to most of my work: operating clusters, helping build hosted Kubernetes, and writing Kubernetes operators. At PlanetScale, that now means running stateful systems like Postgres and MySQL in production. Kubernetes has many faces, but here I want to talk about one face only: why it is so good at running workloads at scale. People ask me what an operator actually does. The canonical answer is: "it reconciles desired state." This is correct, but it also tells you almost nothing. An operator is a feedback controller. It's the same closed loop that runs a thermostat or keeps your car at a fixed speed on cruise control. In our case, the thing being controlled is a database. I have been building these loops for years, and the best way I know to make them click is to ignore Kubernetes at the beginning. Kubernetes is full of control theory, even if we don't call it that in the day-to-day. Before we look at a single line of Kubernetes, we're going to run a production database by hand and slowly let the feedback loop appear on its own. Then we'll map that loop to Kubernetes, with the pieces production needs: a store, watches, queues, retries, and more. At the end, we'll look at what one of these loops looks like in a real operator. https://planetscale.com/blog/the-feedback-loops-behind-kubernetes0,00%
- 11 авг.Client’s GKE Cluster Ate Their Entire VPC GKE pod IP exhaustion is one of the few failure modes that gives you no warning before it goes terminal. I recently stepped into a war room where a client’s primary scaling group had flatlined — workloads cordoned, deployments stuck in Pending, and the estimated cost of the stall nearing $15k per hour in lost transaction volume. The culprit wasn’t traffic. It was a /20 subnet that had quietly run out of address space, and a set of GKE allocation defaults nobody had questioned at design time. The IP Math I Uncovered During Triage: https://www.rack2cloud.com/gke-pod-ip-exhaustion-triage-part-1 The Class E Rescue: https://www.rack2cloud.com/gke-ip-exhaustion-fix-part-20,00%
- 11 авг.🔥 Приглашаем на бесплатный открытый вебинар курса «Observability: мониторинг, логирование, трассировка»: «Системы логирования: ELK, EFK или Graylog?» 🗓 Когда: 17 августа, 20:00 (мск) Логи — один из ключевых источников информации о состоянии системы. Но без правильно выбранного инструмента они превращаются в хаотичный поток данных, в котором сложно найти причину проблемы. На вебинаре сравним популярные системы централизованного логирования и поможем вам выбрать оптимальное решение под вашу инфраструктуру. Что будет на вебинаре: - Чем отличаются ELK, EFK и Graylog и в каких сценариях каждый стек наиболее эффективен - Как устроен процесс сбора, обработки, хранения и поиска логов - Как организовать централизованное логирование для мониторинга и диагностики распределённых систем - На что обратить внимание при выборе системы логирования для своей инфраструктуры В результате вы: - Получите понимание сильных и слабых сторон ELK, EFK и Graylog - Научитесь выбирать подходящее решение под задачи проекта и инфраструктуры - Узнаете лучшие практики построения централизованной системы логирования - Сможете использовать логи для ускорения диагностики и повышения наблюдаемости сервисов Кому будет полезно: DevOps- и SRE-инженерам, системным администраторам, Backend-разработчикам и архитекторам, которым важно быстро находить причины сбоев и анализировать поведение систем. 👉 Зарегистрируйтесь https://vk.cc/d0m7Cs Бесплатное занятие приурочено к старту курса «Observability: мониторинг, логирование, трассировка», на котором вы научитесь строить современные системы наблюдаемости с Prometheus, Grafana, ELK, Tempo и другими инструментами. Реклама. ООО «Отус онлайн-образование», ОГРН 1177746618576, erid: 2VtzqvHEFpX0,00%