OPS/MVS

 View Only

 Extracting Messages from Live OPSLOG with OPSLOG Function

Travis Bream's profile image
Travis Bream posted Aug 10, 2026 04:47 PM

Goal: I need to count the number of message pairs for a "bad" and "good" message over a given amount of time. In this case DFHSM0131 & 2 and DFHSM0133 & 4 over the course of 10 minutes (5 pairs, 10 messages in one CICS region over 10 minutes). The best term I have come up with is "thrashing". Think BAD-GOOD-BAD-GOOD-BAD... I know I can do this using GLVs, but that isn't the most user friendly shall we say and I was looking for a slightly more elegant solution. If thrashing is detected, we need to alert someone.

My first attempt was using the OPSTHRSH function but we encountered a caveat that it is not a rolling count in that it resets its count when the interval expires. This implies it has some sort of internal mechanism that starts the interval when first invoked and then cuts it off after the time has exceeded the interval. Sort of like an elapsed time. Thus, this function will not work for this application since we could be one message away from triggering an alert just to have the interval reset and the counts return to 1.

The next suggestion was to use the OPSLOG function to look back at messages in the LIVE OPSLOG so I set about looking into it and it looked promising but I have hit a stumbling block. I can get the BASIC format (BF) of the function to count correctly but the EXPANDED format (XF) seems to only count the bad messages (TESTBD131 in this case) when the bad message is the one that triggered the rule. When the good message (TESTGD132) is issued, the count of the bad message is 0. For testing, I have a REXX which is doing an ADDRESS WTO with 1 second intervals to generate my test conditions. It issues a bad message, waits 1 second, issues a good message, waits 1 second, issues a bad message and so on. I would use the BF of the function but since my overall goal targets CICS messages, I need to take the job name into account otherwise we could get 5 bad/5 good messages from 5 separate regions and cut a ticket even though there is no issue.

This is the current iteration of my rule:

)MSG TEST*

)PROC

jb_nm = MSG.OJOBNAME or jb_nm = MSG.JOBNAME (I think OJOBNAME may be more correct in this use case but both give the same behavior.)

z = OPSCLEDQ()

var = OPSLOG("EXTRACT TIME(-30) MSGID(TESTBD131) JOBNAME("jb_nm")")

SAY "VAR ="var

RETURN 'NORMAL'

)TERM

)END

I have attempted multiple ways of coding the parameters for the function but this iteration should be all that I need correct? No matter what parameters I seem to use, I always seem to get the same result - a count of 1 for the TESTBD131 message when a TESTBD131 triggers the rule and a count of 0 for the TESTBD131 message when a TESTGD132 triggers the rule.

George Liang's profile image
Broadcom Employee George Liang

Hi, Travis

Please use this keyword format for the JOBNAME:  JOBNAME(name1[,...[,name8]] source(J|O|B))

Then the OPSLOG function is coded like this: 

OPSLOG("EXTRACT TIME(-30) MSGID(TESTBD131)",
       "JOBNAME("jb_nm" SOURCE(O))")

Please let me know if it works.

Thank you

George

Travis Bream's profile image
Travis Bream

George,

I tried that iteration with the JOBNAME and it did the same behavior. Every TESTBD131 would count 1 and TESTGD132 would count TESTBD131 as 0. 

A colleague and I did get SCANTEXT to work but it kept picking up on the OPS1370O message coming from the REXX that was generating the messages and thus duplicating the count. We attempted to filter out that message but were unsuccessful. I have some more testing to do as well.

George Liang's profile image
Broadcom Employee George Liang

Hi, Travis

Could you please share the code of your OPSLOG function call ?  Did you add the  SOURCE(O) into the JOBNAME keyword?

Thanks

George

Travis Bream's profile image
Travis Bream

Apologies for the delay in response, other work obligations took priority. I coded the OPSLOG function exactly as you suggessted and it exhibited the behavior I saw before. Below I have pasted my code as well as the output so you can see the count changing each time.

RULE CODE

)MSG TEST*                                          
)INIT                                               
)PROC                                               
diag_pfx = "DIAG:" OPSINFO('PROGRAM')               
jb_nm = MSG.OJOBNAME                                
SAY diag_pfx "*** BEGIN OPSLOG() TEST ***"          
SAY diag_pfx "JB_NM       ="jb_nm                   
                                                    
z = OPSCLEDQ()                                      
test1 = OPSLOG("EXTRACT TIME(-30) MSGID(TESTBD131)",
        "JOBNAME("jb_nm" SOURCE(O))")               
                                                    
SAY diag_pfx "TEST1       ="test1                   
                                                    
RETURN 'NORMAL'                                     

OUTPUT

Date  TIME     ----+----1----+----2----+----3----+----4----+----5----+----6----
21AUG 13:30:14 SOSTEST 4                                                       <== Command to initiate my test REXX.
21AUG 13:30:14 OPS3724O TSO TOTMB.SOSTEST Sent CMD=OX P('OEMDS.OPSMVS.RULES.TOT<== Test REXX execting...
21AUG 13:30:14 OPS3092O OX P('OEMDS.OPSMVS.RULES.TOTMB.MF1T(SOSTSTX)') A(4)    <== ...

21AUG 13:30:14 OPS1370O OPSOSF   X'0000' X'0000' X'0000' NONE  300 TESTBD131  a<== REXX issues WTO.
21AUG 13:30:14 TESTBD131 applid CICS is under stress (short on storage below 16<== Test BAD WTO. Pass 1.
21AUG 13:30:14 OPS1000O DIAG: TOTMB.OPSLGTST *** BEGIN OPSLOG() TEST ***       <== Diagnostic message from my rule.
21AUG 13:30:14 OPS1000O DIAG: TOTMB.OPSLGTST JB_NM       =OPSOSF               <== Diagnostic message from my rule. Jobname from RULE.
21AUG 13:30:14 OPS1000O DIAG: TOTMB.OPSLGTST TEST1       =1                    <== Diagnostic message from my rule. Count from OPSLOG().
21AUG 13:30:15 OPS1370O OPSOSF   X'0000' X'0000' X'0000' NONE  300 TESTGD132  a
21AUG 13:30:15 TESTGD132 applid CICS is no longer short on storage below 16MB. <== Test GOOD WTO. Pass 1.
21AUG 13:30:15 OPS1000O DIAG: TOTMB.OPSLGTST *** BEGIN OPSLOG() TEST ***       
21AUG 13:30:15 OPS1000O DIAG: TOTMB.OPSLGTST JB_NM       =OPSOSF               <== Diagnostic message from my rule. Jobname from RULE.
21AUG 13:30:15 OPS1000O DIAG: TOTMB.OPSLGTST TEST1       =0                    <== Diagnostic message from my rule. Count from OPSLOG().
21AUG 13:30:16 OPS1370O OPSOSF   X'0000' X'0000' X'0000' NONE  300 TESTBD131  a
21AUG 13:30:16 TESTBD131 applid CICS is under stress (short on storage below 16<== Test BAD WTO. Pass 2.
21AUG 13:30:16 OPS1000O DIAG: TOTMB.OPSLGTST *** BEGIN OPSLOG() TEST ***       
21AUG 13:30:16 OPS1000O DIAG: TOTMB.OPSLGTST JB_NM       =OPSOSF               <== Diagnostic message from my rule. Jobname from RULE.
21AUG 13:30:16 OPS1000O DIAG: TOTMB.OPSLGTST TEST1       =1                    <== Diagnostic message from my rule. Count from OPSLOG().
21AUG 13:30:17 OPS1370O OPSOSF   X'0000' X'0000' X'0000' NONE  300 TESTGD132  a
21AUG 13:30:17 TESTGD132 applid CICS is no longer short on storage below 16MB. <== Test GOOD WTO. Pass 2.
21AUG 13:30:17 OPS1000O DIAG: TOTMB.OPSLGTST *** BEGIN OPSLOG() TEST ***       
21AUG 13:30:17 OPS1000O DIAG: TOTMB.OPSLGTST JB_NM       =OPSOSF               <== Diagnostic message from my rule. Jobname from RULE.
21AUG 13:30:17 OPS1000O DIAG: TOTMB.OPSLGTST TEST1       =0                    <== Diagnostic message from my rule. Count from OPSLOG().
21AUG 13:30:18 OPS1370O OPSOSF   X'0000' X'0000' X'0000' NONE  300 TESTBD131  a
21AUG 13:30:18 TESTBD131 applid CICS is under stress (short on storage below 16<== Test BAD WTO. Pass 2.
21AUG 13:30:18 OPS1000O DIAG: TOTMB.OPSLGTST *** BEGIN OPSLOG() TEST ***       
21AUG 13:30:18 OPS1000O DIAG: TOTMB.OPSLGTST JB_NM       =OPSOSF               <== Diagnostic message from my rule. Jobname from RULE.
21AUG 13:30:18 OPS1000O DIAG: TOTMB.OPSLGTST TEST1       =1                    <== Diagnostic message from my rule. Count from OPSLOG().
21AUG 13:30:19 OPS1370O OPSOSF   X'0000' X'0000' X'0000' NONE  300 TESTGD132  a
21AUG 13:30:19 TESTGD132 applid CICS is no longer short on storage below 16MB. <== Test GOOD WTO. Pass 2.
21AUG 13:30:19 OPS1000O DIAG: TOTMB.OPSLGTST *** BEGIN OPSLOG() TEST ***       
21AUG 13:30:19 OPS1000O DIAG: TOTMB.OPSLGTST JB_NM       =OPSOSF               <== Diagnostic message from my rule. Jobname from RULE.
21AUG 13:30:19 OPS1000O DIAG: TOTMB.OPSLGTST TEST1       =0                    <== Diagnostic message from my rule. Count from OPSLOG().
21AUG 13:30:20 OPS1370O OPSOSF   X'0000' X'0000' X'0000' NONE  300 TESTBD131  a
21AUG 13:30:20 TESTBD131 applid CICS is under stress (short on storage below 16<== Test BAD WTO. Pass 2.
21AUG 13:30:20 OPS1000O DIAG: TOTMB.OPSLGTST *** BEGIN OPSLOG() TEST ***       
21AUG 13:30:20 OPS1000O DIAG: TOTMB.OPSLGTST JB_NM       =OPSOSF               <== Diagnostic message from my rule. Jobname from RULE.
21AUG 13:30:20 OPS1000O DIAG: TOTMB.OPSLGTST TEST1       =1                    <== Diagnostic message from my rule. Count from OPSLOG().
21AUG 13:30:21 OPS1370O OPSOSF   X'0000' X'0000' X'0000' NONE  300 TESTGD132  a
21AUG 13:30:21 TESTGD132 applid CICS is no longer short on storage below 16MB. <== Test GOOD WTO. Pass 2.
21AUG 13:30:21 OPS1000O DIAG: TOTMB.OPSLGTST *** BEGIN OPSLOG() TEST ***       
21AUG 13:30:21 OPS1000O DIAG: TOTMB.OPSLGTST JB_NM       =OPSOSF               <== Diagnostic message from my rule. Jobname from RULE.
21AUG 13:30:21 OPS1000O DIAG: TOTMB.OPSLGTST TEST1       =0                    <== Diagnostic message from my rule. Count from OPSLOG().
21AUG 13:30:22 OPS1370O OPSOSF   X'0000' X'0000' X'0000' NONE  300 TESTBD131  a
21AUG 13:30:22 TESTBD131 applid CICS is under stress (short on storage below 16<== Test BAD WTO. Pass 2.
21AUG 13:30:22 OPS1000O DIAG: TOTMB.OPSLGTST *** BEGIN OPSLOG() TEST ***       
21AUG 13:30:22 OPS1000O DIAG: TOTMB.OPSLGTST JB_NM       =OPSOSF               <== Diagnostic message from my rule. Jobname from RULE.
21AUG 13:30:22 OPS1000O DIAG: TOTMB.OPSLGTST TEST1       =1                    <== Diagnostic message from my rule. Count from OPSLOG().
21AUG 13:30:23 OPS1370O OPSOSF   X'0000' X'0000' X'0000' NONE  300 TESTGD132  a
21AUG 13:30:23 TESTGD132 applid CICS is no longer short on storage below 16MB. <== Test GOOD WTO. Pass 2.
21AUG 13:30:23 OPS1000O DIAG: TOTMB.OPSLGTST *** BEGIN OPSLOG() TEST ***       
21AUG 13:30:23 OPS1000O DIAG: TOTMB.OPSLGTST JB_NM       =OPSOSF               <== Diagnostic message from my rule. Jobname from RULE.
21AUG 13:30:23 OPS1000O DIAG: TOTMB.OPSLGTST TEST1       =0                    <== Diagnostic message from my rule. Count from OPSLOG().
21AUG 13:30:24 OPS3092O READY                                                  <== END REXX EXECUTION.