How Do You Capture & Revisit Your Investigations? (Where Does It Break?)
Summary: Ian Marge, a member of the LogicMonitor product team, seeks insight from the community about how teams handle incident investigations and preserve related information. They aim to understand how investigations are currently captured, what challenges exist in preserving investigation context, and how often past investigations are revisited. They also show interest in learning what tools are used and where improvements are necessary for better collaboration and trust in past analyses. Feedback from the community is encouraged to assist in enhancing LogicMonitor's investigation workflows.
Hi LogicMonitor Community,
I’m part of the LogicMonitor product team, and we’re researching how teams investigate incidents and just as importantly, how that work gets preserved (or lost) over time.
In many environments, investigation context ends up spread across dashboards, queries, tickets, screenshots, and docs.
We’re trying to better understand how this works for you today and where it breaks down.
When you're investigating incidents or performance issues:
How do you capture what you reviewed (queries, dashboards, alerts, time ranges) so others can understand or revisit your work later?
How often do you need to revisit or re-run past investigations, and how easy or painful is that process today?
Where does context typically get lost during or after an investigation?
What tools or workflows do you rely on today, and what don’t they do well?
What makes it difficult to trust or reuse past analysis when similar issues come up again?
We’re also exploring a concept around making investigations easier to revisit:
If you could go back and see an investigation exactly as it was (same queries, time ranges, and context), how would that change your workflow?
Your feedback will directly shape how we think about improving investigation and collaboration workflows in LogicMonitor.
Even quick examples or high-level thoughts are incredibly helpful.
Feel free to reply here or message me directly if you prefer.
Thanks in advance for helping us build something truly useful for this community.
Best,
Ian
Michael Dieter
·3 months agoHey Ian
Currently, we wind up with a blob-shaped sprawl of attempts to surface contextual evidence that might even be spread across multiple responders' devices:
several tabs open to different Logicmonitor resources, logs, websites and/or dashboards; a local laptop command prompt; a local laptop ssh console; a Palo Alto Panorama tab; a Juniper Mist tab, etc. Of course in the background chat and/or voice calls are going on using Teams. This is all completely ephemeral, with effectively nothing preserved for future replay/reference save a few possible notes that might make it into a ticket resolution (if a ticket even existed).
I've sometimes experimented in the past with trying to add extensive notes to Alert Acknowledgements and add verbose Ops Notes but I've never found a sustainable way to preserve the ability to later recall/search and there seems to be less and less demand for recall in our environment beyond a "TLDR".
Its probably fair to tie this all to the (mostly tacit) demands, SLAs and SLEs that our organization chooses to live with. But your concept is much more interesting to orgs that have a different way of valuing response and service restoration